system
A voice-controlled system simplifies video shooting, editing, and uploading using AI, allowing users to easily create and share high-quality content.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
The process from video shooting to editing and uploading is time-consuming, particularly difficult for beginners.
A system comprising a reception unit, shooting unit, and upload unit that uses voice commands to facilitate easy video recording, editing, and uploading, utilizing AI for analysis and automation.
Enables users to easily create and share high-quality videos by simplifying the process through voice-controlled video shooting, editing, and uploading.
Smart Images

Figure 2026072861000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there is a problem that the process from video shooting to editing and uploading is time-consuming and particularly difficult for beginners.
[0005] The system according to the embodiment aims to easily perform video shooting, editing, and uploading using voice commands.
Means for Solving the Problems
[0006] The system according to this embodiment comprises a reception unit, a shooting unit, an editing unit, and an upload unit. The reception unit receives voice commands. The shooting unit performs shooting based on the voice commands received by the reception unit. The editing unit analyzes the data shot by the shooting unit and performs optimal editing. The upload unit automatically uploads the video edited by the editing unit to a video site. [Effects of the Invention]
[0007] The system according to this embodiment allows for easy recording, editing, and uploading of videos using voice commands. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The AI glasses system according to an embodiment of the present invention is a system that allows users to edit videos and add subtitles simply by wearing the glasses and filming. This AI glasses system can receive filming instructions via voice commands, automatically edit the filmed video, add comments, and upload it to a video site. For example, a user puts on the AI glasses and inputs a voice command such as "Film the cooking." The AI glasses system then starts filming according to the instructions, and the AI editing software analyzes the filmed video and performs optimal editing. For example, it might highlight important scenes or cut out unnecessary parts to make the cooking process easier to understand. The AI also automatically adds comments to create a video that is easy for viewers to understand. Furthermore, the AI glasses system automatically uploads the edited video to a video site. This allows users to easily create and share high-quality videos without performing complex editing tasks. The development of this AI glasses system and AI editing software will enable anyone to easily become a video streamer, bringing happiness to people through a filming revolution. Thus, the AI glasses system allows users to easily create and share high-quality videos.
[0029] The AI glasses system according to this embodiment comprises a reception unit, a shooting unit, an editing unit, and an upload unit. The reception unit receives voice commands. Voice commands include, but are not limited to, examples such as "Start shooting" or "Shoot the food." The reception unit analyzes the voice commands using, for example, speech recognition technology and generates appropriate instructions. The shooting unit takes pictures based on the voice commands received by the reception unit. The shooting unit has, for example, a built-in camera and shoots images from the user's point of view. The shooting unit has, for example, an image stabilization function and can shoot stable images. The shooting unit can also shoot wide-angle images using, for example, a wide-angle lens. The editing unit analyzes the data captured by the shooting unit and performs optimal editing. The editing unit uses, for example, AI to analyze the captured data and highlight important scenes or cut out unnecessary parts. The editing unit uses, for example, AI to automatically add comments to create a video that is easy for viewers to understand. The editing unit can also use, for example, AI to add music or sound effects to enhance the appeal of the video. The upload unit automatically uploads videos edited by the editorial unit to a video site. The upload unit, for example, accesses the video site via an internet connection and uploads the video. The upload unit also automatically generates video titles and descriptions, for example, to provide engaging content for viewers. This allows the AI glasses system according to the embodiment to enable users to easily create and share high-quality videos.
[0030] The reception unit receives voice commands. Voice commands include, but are not limited to, examples such as "Start shooting" or "Take a picture of the food." The reception unit analyzes voice commands using, for example, speech recognition technology and generates appropriate instructions. Specifically, the speech recognition technology converts the user's voice into a digital signal and matches it with a voice database to identify the content of the command. The speech recognition technology incorporates noise cancellation to remove ambient noise and achieve accurate speech recognition. Furthermore, the speech recognition technology learns the characteristics of the user's voice and provides recognition accuracy optimized for each individual user. For example, if a user says "Start shooting," the reception unit analyzes this command and sends an instruction to the shooting unit to start shooting. In addition, to increase the variety of voice commands, the reception unit uses natural language processing technology to be able to handle different expressions and phrasing. For example, it can similarly recognize commands such as "Take a video" or "Start recording" and generate appropriate instructions. This allows the reception unit to accurately and quickly analyze user voice commands and improve the overall usability of the system.
[0031] The shooting unit performs shooting based on voice commands received by the reception unit. The shooting unit, for example, has a built-in camera that captures video from the user's point of view. Specifically, the camera is equipped with a high-resolution sensor, enabling it to capture clear images. Furthermore, the shooting unit has an image stabilization function, allowing it to capture stable video even when the user is moving. The image stabilization function uses a gyro sensor and accelerometer to detect camera movement and correct it in real time. The shooting unit can also capture a wide range of images using a wide-angle lens. The wide-angle lens expands the user's field of view, allowing it to capture more information at once. In addition, the shooting unit is equipped with a night mode and a high-sensitivity sensor to capture high-quality video even in low-light environments. As a result, the shooting unit can adapt to various environments and situations and always provide high-quality video. For example, if the user issues the command "Take a picture of the food," the shooting unit will automatically activate the camera and capture video of the food from the user's point of view. This allows the shooting unit to perform shooting quickly and accurately according to the user's instructions, improving the overall usability of the system.
[0032] The editorial team analyzes the data captured by the filming team and performs optimal editing. For example, the editorial team uses AI to analyze the footage, highlighting important scenes and cutting out unnecessary parts. Specifically, AI uses video analysis technology to evaluate the content and importance of scenes, creating videos that are engaging for viewers. For example, AI can use facial recognition technology to analyze people's expressions and movements, capturing changes in emotion. AI can also use voice analysis technology to analyze the content and tone of conversations, highlighting important statements and emotional moments. Furthermore, AI can automatically add comments to make the videos easier for viewers to understand. For example, it can display explanations of ingredients and cooking procedures as text for cooking videos. In addition, AI can add music and sound effects to enhance the appeal of the videos. Music and sound effects play a role in emphasizing the atmosphere and emotions of a scene and attracting the viewer's interest. This allows the editorial team to quickly and effectively edit the footage and provide engaging content for viewers. For example, the editorial team can analyze cooking videos filmed by users, highlight important scenes, cut out unnecessary parts, and add comments and music to create cooking videos that are easy to understand and appealing to viewers.
[0033] The upload unit automatically uploads videos edited by the editorial unit to video sharing sites. For example, the upload unit accesses video sharing sites via an internet connection and uploads the videos. Specifically, the upload unit uses the video sharing site's API to send video files to the server and handles publication settings and tagging. Furthermore, the upload unit automatically generates video titles and descriptions, providing engaging content for viewers. Natural language generation technology is used to automatically generate titles and descriptions, creating appropriate text based on the video's content. For example, for a cooking video, the title might be something specific and interesting like "Easy Recipe: How to Make Delicious Pasta," and the description might provide detailed information such as, "This video shows how to make delicious pasta. We explain the ingredients and steps in detail, so please take a look." The upload unit also automatically sets the video's publication range and privacy settings, allowing users to choose a publication method that aligns with their intentions. This enables the upload unit to easily share high-quality videos without requiring any effort from the user. For example, a cooking video filmed and edited by a user can be automatically uploaded to a video sharing site by the upload unit, which then sets an appropriate title and description, providing engaging content for viewers.
[0034] The reception unit can interpret voice commands optimally by referring to the user's past command history. For example, the reception unit can prioritize interpreting commands that the user has frequently used in the past. For example, the reception unit can predict and interpret commands used in specific situations based on the user's past command history. For example, the reception unit can analyze the user's past command history and correct commands that are easily misunderstood. This makes it possible to interpret commands optimally for the user by referring to past command history. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can input the user's past command history data into a generating AI and have the generating AI perform the optimal command interpretation.
[0035] The reception unit can filter out ambient noise and remove noise when receiving voice commands. For example, if the ambient noise is loud, the reception unit can use AI to clear the voice command using noise cancellation technology. For example, if there is strong wind noise, the reception unit can use AI to remove wind noise and recognize the voice command. For example, if background music is playing, the reception unit can use AI to filter the music and accurately recognize the voice command. In this way, by filtering out ambient noise, voice commands can be accurately recognized without being affected by noise. Some or all of the above processing in the reception unit may be performed using AI, or not using AI. For example, the reception unit can input ambient noise data into a generating AI and have the generating AI perform noise reduction.
[0036] The reception unit can prioritize receiving commands that are highly relevant when receiving voice commands, taking into account the user's geographical location. For example, if the user is in a specific location, the reception unit can prioritize receiving commands related to that location. For example, if the user is on the move, the reception unit can prioritize receiving commands related to movement. For example, if the user is at home, the reception unit can prioritize receiving commands related to activities at home. In this way, by taking into account the user's geographical location, the reception unit can prioritize receiving commands that are highly relevant. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can input the user's geographical location data into a generating AI and have the generating AI perform the priority reception of highly relevant commands.
[0037] The reception unit can analyze the user's social media activity when receiving a voice command and suggest relevant commands. For example, if the user has posted about a specific topic on social media, the reception unit can suggest commands related to that topic. For example, if the user has participated in a specific event on social media, the reception unit can suggest commands related to that event. For example, if the user has checked in to a specific location on social media, the reception unit can suggest commands related to that location. In this way, by analyzing social media activity, commands relevant to the user can be suggested. Some or all of the above processing in the reception unit may be performed using AI, for example, or not using AI. For example, the reception unit can input the user's social media activity data into a generating AI and have the generating AI execute the suggestion of relevant commands.
[0038] The camera unit can predict the subject's movement and take a picture at the optimal timing. For example, the camera unit can predict the direction the subject will move and release the shutter at the optimal time. For example, the camera unit can predict the subject's speed and adjust the shutter speed to prevent blurring. For example, the camera unit can predict the subject's movement and automatically switch to continuous shooting mode. This allows the camera unit to take a picture at the optimal timing by predicting the subject's movement. Some or all of the above processes in the camera unit may be performed using AI, for example, or without AI. For example, the camera unit can input subject movement data into a generating AI and have the generating AI execute a shot at the optimal timing.
[0039] The shooting unit can automatically adjust the ambient light level and color temperature during shooting to acquire optimal images. For example, if the ambient light level is insufficient, the AI in the shooting unit can automatically increase the ISO sensitivity. For example, if the color temperature is inappropriate, the AI in the shooting unit can automatically adjust the white balance. For example, if there are extreme variations in light intensity, the AI in the shooting unit can automatically adjust the exposure. In this way, optimal images can be acquired by automatically adjusting the ambient light level and color temperature. Some or all of the above processing in the shooting unit may be performed using AI, or not. For example, the shooting unit can input ambient light level and color temperature data into a generating AI and have the generating AI perform the acquisition of optimal images.
[0040] The shooting unit can automatically select the optimal shooting settings while considering the user's geographical location information. For example, if the user is outdoors, the AI can automatically adjust the exposure and white balance. For example, if the user is indoors, the AI can automatically adjust the ISO sensitivity and shutter speed. For example, if the user is at a specific tourist destination, the shooting unit can automatically select shooting settings appropriate for that location. In this way, the optimal shooting settings can be automatically selected by considering the user's geographical location information. Some or all of the above processing in the shooting unit may be performed using AI, for example, or without AI. For example, the shooting unit can input the user's geographical location information data into a generating AI and have the generating AI perform the automatic selection of the optimal shooting settings.
[0041] The shooting unit can analyze the user's social media activity during shooting and suggest relevant shooting scenes. For example, if the user posts about a specific topic on social media, the shooting unit can suggest shooting scenes related to that topic. For example, if the user participates in a specific event on social media, the shooting unit can suggest shooting scenes related to that event. For example, if the user checks in to a specific location on social media, the shooting unit can suggest shooting scenes related to that location. In this way, by analyzing social media activity, it is possible to suggest shooting scenes relevant to the user. Some or all of the above processing in the shooting unit may be performed using AI, for example, or not using AI. For example, the shooting unit can input the user's social media activity data into a generating AI and have the generating AI suggest relevant shooting scenes.
[0042] The editorial department can analyze the content of the filmed data during editing and automatically highlight important scenes. For example, in the case of a cooking video, the editorial department can highlight important cooking steps. For example, in the case of an event video, the editorial department can highlight highlight scenes. For example, in the case of an interview video, the editorial department can highlight important statements. In this way, important scenes can be automatically highlighted by analyzing the content of the filmed data. Some or all of the above processing in the editorial department may be performed using AI, for example, or without AI. For example, the editorial department can input the filmed data into a generating AI and have the generating AI perform the highlighting of important scenes.
[0043] The editorial team can automatically cut out unnecessary scenes during editing to create smoother footage. For example, in cooking videos, the editorial team can cut out waiting times and scenes of mistakes. For example, in event videos, the editorial team can cut out unnecessary scenes. For example, in interview videos, the editorial team can cut out unnecessary remarks and silences. This allows for smoother footage by automatically cutting out unnecessary scenes. Some or all of the above processing in the editorial team may be performed using AI, for example, or not. For example, the editorial team can input the shooting data into a generating AI and have the generating AI cut out unnecessary scenes.
[0044] The editorial team can prioritize editing scenes that are highly relevant, taking into account the user's geographical location information. For example, the editorial team can prioritize editing scenes that the user took at a specific location. For example, the editorial team can prioritize editing scenes that the user took while traveling. For example, the editorial team can prioritize editing scenes that the user took at an event venue. This allows for the prioritization of highly relevant scenes by considering the user's geographical location information. Some or all of the above processing by the editorial team may be performed using AI, for example, or not using AI. For example, the editorial team can input the user's geographical location data into a generating AI and have the generating AI perform the priority editing of highly relevant scenes.
[0045] The editorial team can analyze users' social media activity during editing and automatically add relevant comments and effects. For example, if a user posts about a specific topic on social media, the editorial team can automatically add comments related to that topic. For example, if a user participates in a specific event on social media, the editorial team can automatically add effects related to that event. For example, if a user checks in to a specific location on social media, the editorial team can automatically add comments and effects related to that location. This allows for the automatic addition of relevant comments and effects by analyzing social media activity. Some or all of the above processing by the editorial team may be performed using AI, for example, or not. For example, the editorial team can input user social media activity data into a generating AI and have the generating AI perform the automatic addition of relevant comments and effects.
[0046] The upload unit can automatically generate video metadata and perform optimal tagging during upload. For example, the upload unit can analyze the video content and automatically tag it with relevant keywords. For example, the upload unit can automatically generate different tags for each scene in the video, making it easier for viewers to search. For example, the upload unit can automatically generate optimal tags based on the video title and description. This enables optimal tagging by automatically generating video metadata. Some or all of the above processes in the upload unit may be performed using AI, for example, or without AI. For example, the upload unit can input video content data into a generation AI and have the generation AI perform automatic metadata generation and tagging.
[0047] The upload unit can automatically adjust the video's privacy settings during upload to set the optimal visibility. For example, the upload unit can automatically select the optimal privacy settings by referring to the user's past visibility settings. For example, the upload unit can automatically adjust the visibility depending on the video's content. For example, the upload unit can analyze the user's social media activity and suggest the optimal visibility. This allows the optimal visibility to be set by automatically adjusting the video's privacy settings. Some or all of the above processes in the upload unit may be performed using AI, for example, or without AI. For example, the upload unit can input video content data into a generating AI and have the generating AI perform the automatic adjustment of privacy settings.
[0048] The upload unit can select the most suitable video site by considering the user's geographical location information during the upload process. For example, if the user is in a specific region, the upload unit can select a video site popular in that region. For example, if the user is traveling, the upload unit can select a video site popular in the travel destination. For example, if the user is participating in a specific event, the upload unit can select a video site related to that event. In this way, the most suitable video site can be selected by considering the user's geographical location information. Some or all of the above processing in the upload unit may be performed using AI, for example, or without AI. For example, the upload unit can input the user's geographical location data into a generating AI and have the generating AI select the most suitable video site.
[0049] The upload unit can analyze a user's social media activity during upload and automatically share it to relevant platforms. For example, if a user posts about a specific topic on social media, the upload unit can automatically share videos related to that topic. For example, if a user participates in a specific event on social media, the upload unit can automatically share videos related to that event. For example, if a user checks in to a specific location on social media, the upload unit can automatically share videos related to that location. This allows for automatic sharing to relevant platforms by analyzing social media activity. Some or all of the above processing in the upload unit may be performed using AI, for example, or without AI. For example, the upload unit can input the user's social media activity data into a generating AI and have the generating AI perform automatic sharing to relevant platforms.
[0050] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0051] The AI glasses system can also be equipped with a health management unit that monitors the user's health status. This unit can, for example, measure the user's heart rate and blood pressure and issue a warning if an abnormality is detected. It can also, for example, record the user's steps and calories burned to support daily health management. Furthermore, it can analyze the user's sleep patterns and provide advice to promote high-quality sleep. This allows the AI glasses system to monitor the user's health status in real time and support health management.
[0052] AI glasses systems can also include a recommendation system that suggests content based on the user's hobbies and interests. For example, if the user likes music, the recommendation system can provide the latest music trends and recommended playlists. If the user likes movies, the recommendation system can provide the latest movie information and recommended movies. If the user likes sports, the recommendation system can provide the latest match results and recommended sporting events. This can enhance the user's entertainment experience by recommending content based on their hobbies and interests.
[0053] The AI glasses system can also be equipped with a learning support unit to assist the user's learning. For example, if the user is learning a new language, the learning support unit can provide advice on pronunciation and grammar. If the user is studying for an exam, the learning support unit can provide efficient learning methods and key points. If the user wants to improve a hobby skill, the learning support unit can provide relevant learning materials and practice methods. This supports the user's learning, thereby promoting skill improvement and knowledge acquisition.
[0054] The AI glasses system can also be equipped with a schedule management unit to support the user's schedule management. For example, the schedule management unit can automatically organize the user's schedule and remind them of important events and tasks. For example, the schedule management unit can suggest the optimal time allocation based on the user's schedule. For example, the schedule management unit can provide traffic and weather information according to the user's schedule. This allows for efficient time management by supporting the user's schedule management.
[0055] The AI glasses system can also be equipped with a travel support unit to assist users during their travels. This unit could, for example, recommend tourist attractions and restaurants based on the user's current location. It could also suggest optimal routes based on the user's travel schedule. Furthermore, it could provide customized travel plans tailored to the user's preferences. This would allow the system to support users' travels and provide a more fulfilling travel experience.
[0056] The following briefly describes the processing flow for example form 1.
[0057] Step 1: The reception desk receives voice commands. Voice commands include phrases such as "Start shooting" or "Take a picture of the food." The reception desk uses voice recognition technology to analyze the voice commands and generate appropriate instructions. Step 2: The shooting unit takes pictures based on the voice commands received by the reception unit. The shooting unit has a built-in camera and captures images from the user's point of view. It can capture stable, wide-angle images using image stabilization and a wide-angle lens. Step 3: The editorial team analyzes the footage shot by the camera crew and performs optimal editing. Using AI, they analyze the footage to highlight important scenes and cut out unnecessary parts. Furthermore, the AI automatically adds commentary, music, and sound effects to create a video that is easy to understand and engaging for viewers. Step 4: The uploading unit automatically uploads the videos edited by the editorial team to the video site. It accesses the video site via an internet connection and uploads the videos. Furthermore, it automatically generates video titles and descriptions to provide engaging content for viewers.
[0058] (Example of form 2) The AI glasses system according to an embodiment of the present invention is a system that allows users to edit videos and add subtitles simply by wearing the glasses and filming. This AI glasses system can receive filming instructions via voice commands, automatically edit the filmed video, add comments, and upload it to a video site. For example, a user puts on the AI glasses and inputs a voice command such as "Film the cooking." The AI glasses system then starts filming according to the instructions, and the AI editing software analyzes the filmed video and performs optimal editing. For example, it might highlight important scenes or cut out unnecessary parts to make the cooking process easier to understand. The AI also automatically adds comments to create a video that is easy for viewers to understand. Furthermore, the AI glasses system automatically uploads the edited video to a video site. This allows users to easily create and share high-quality videos without performing complex editing tasks. The development of this AI glasses system and AI editing software will enable anyone to easily become a video streamer, bringing happiness to people through a filming revolution. Thus, the AI glasses system allows users to easily create and share high-quality videos.
[0059] The AI glasses system according to this embodiment comprises a reception unit, a shooting unit, an editing unit, and an upload unit. The reception unit receives voice commands. Voice commands include, but are not limited to, examples such as "Start shooting" or "Shoot the food." The reception unit analyzes the voice commands using, for example, speech recognition technology and generates appropriate instructions. The shooting unit takes pictures based on the voice commands received by the reception unit. The shooting unit has, for example, a built-in camera and shoots images from the user's point of view. The shooting unit has, for example, an image stabilization function and can shoot stable images. The shooting unit can also shoot wide-angle images using, for example, a wide-angle lens. The editing unit analyzes the data captured by the shooting unit and performs optimal editing. The editing unit uses, for example, AI to analyze the captured data and highlight important scenes or cut out unnecessary parts. The editing unit uses, for example, AI to automatically add comments to create a video that is easy for viewers to understand. The editing unit can also use, for example, AI to add music or sound effects to enhance the appeal of the video. The upload unit automatically uploads videos edited by the editorial unit to a video site. The upload unit, for example, accesses the video site via an internet connection and uploads the video. The upload unit also automatically generates video titles and descriptions, for example, to provide engaging content for viewers. This allows the AI glasses system according to the embodiment to enable users to easily create and share high-quality videos.
[0060] The reception unit receives voice commands. Voice commands include, but are not limited to, examples such as "Start shooting" or "Take a picture of the food." The reception unit analyzes voice commands using, for example, speech recognition technology and generates appropriate instructions. Specifically, the speech recognition technology converts the user's voice into a digital signal and matches it with a voice database to identify the content of the command. The speech recognition technology incorporates noise cancellation to remove ambient noise and achieve accurate speech recognition. Furthermore, the speech recognition technology learns the characteristics of the user's voice and provides recognition accuracy optimized for each individual user. For example, if a user says "Start shooting," the reception unit analyzes this command and sends an instruction to the shooting unit to start shooting. In addition, to increase the variety of voice commands, the reception unit uses natural language processing technology to be able to handle different expressions and phrasing. For example, it can similarly recognize commands such as "Take a video" or "Start recording" and generate appropriate instructions. This allows the reception unit to accurately and quickly analyze user voice commands and improve the overall usability of the system.
[0061] The shooting unit performs shooting based on voice commands received by the reception unit. The shooting unit, for example, has a built-in camera that captures video from the user's point of view. Specifically, the camera is equipped with a high-resolution sensor, enabling it to capture clear images. Furthermore, the shooting unit has an image stabilization function, allowing it to capture stable video even when the user is moving. The image stabilization function uses a gyro sensor and accelerometer to detect camera movement and correct it in real time. The shooting unit can also capture a wide range of images using a wide-angle lens. The wide-angle lens expands the user's field of view, allowing it to capture more information at once. In addition, the shooting unit is equipped with a night mode and a high-sensitivity sensor to capture high-quality video even in low-light environments. As a result, the shooting unit can adapt to various environments and situations and always provide high-quality video. For example, if the user issues the command "Take a picture of the food," the shooting unit will automatically activate the camera and capture video of the food from the user's point of view. This allows the shooting unit to perform shooting quickly and accurately according to the user's instructions, improving the overall usability of the system.
[0062] The editorial team analyzes the data captured by the filming team and performs optimal editing. For example, the editorial team uses AI to analyze the footage, highlighting important scenes and cutting out unnecessary parts. Specifically, AI uses video analysis technology to evaluate the content and importance of scenes, creating videos that are engaging for viewers. For example, AI can use facial recognition technology to analyze people's expressions and movements, capturing changes in emotion. AI can also use voice analysis technology to analyze the content and tone of conversations, highlighting important statements and emotional moments. Furthermore, AI can automatically add comments to make the videos easier for viewers to understand. For example, it can display explanations of ingredients and cooking procedures as text for cooking videos. In addition, AI can add music and sound effects to enhance the appeal of the videos. Music and sound effects play a role in emphasizing the atmosphere and emotions of a scene and attracting the viewer's interest. This allows the editorial team to quickly and effectively edit the footage and provide engaging content for viewers. For example, the editorial team can analyze cooking videos filmed by users, highlight important scenes, cut out unnecessary parts, and add comments and music to create cooking videos that are easy to understand and appealing to viewers.
[0063] The upload unit automatically uploads videos edited by the editorial unit to video sharing sites. For example, the upload unit accesses video sharing sites via an internet connection and uploads the videos. Specifically, the upload unit uses the video sharing site's API to send video files to the server and handles publication settings and tagging. Furthermore, the upload unit automatically generates video titles and descriptions, providing engaging content for viewers. Natural language generation technology is used to automatically generate titles and descriptions, creating appropriate text based on the video's content. For example, for a cooking video, the title might be something specific and interesting like "Easy Recipe: How to Make Delicious Pasta," and the description might provide detailed information such as, "This video shows how to make delicious pasta. We explain the ingredients and steps in detail, so please take a look." The upload unit also automatically sets the video's publication range and privacy settings, allowing users to choose a publication method that aligns with their intentions. This enables the upload unit to easily share high-quality videos without requiring any effort from the user. For example, a cooking video filmed and edited by a user can be automatically uploaded to a video sharing site by the upload unit, which then sets an appropriate title and description, providing engaging content for viewers.
[0064] The reception unit can estimate the user's emotions and adjust the accuracy of voice command recognition based on the estimated emotions. For example, if the user is nervous, the reception unit can use AI to adjust the tone and speed of the voice to improve the accuracy of voice command recognition. For example, if the user is relaxed, the reception unit can use AI to return the accuracy of voice command recognition to normal settings. For example, if the user is in a hurry, the reception unit can use AI to quickly adjust the accuracy of voice command recognition and respond immediately. This allows for more accurate recognition of voice commands by adjusting the accuracy of voice command recognition according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception unit may be performed using AI or not using AI. For example, the reception unit can input the user's voice data into the generative AI and have the generative AI perform the adjustment of the accuracy of voice command recognition.
[0065] The reception unit can interpret voice commands optimally by referring to the user's past command history. For example, the reception unit can prioritize interpreting commands that the user has frequently used in the past. For example, the reception unit can predict and interpret commands used in specific situations based on the user's past command history. For example, the reception unit can analyze the user's past command history and correct commands that are easily misunderstood. This makes it possible to interpret commands optimally for the user by referring to past command history. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can input the user's past command history data into a generating AI and have the generating AI perform the optimal command interpretation.
[0066] The reception unit can filter out ambient noise and remove noise when receiving voice commands. For example, if the ambient noise is loud, the reception unit can use AI to clear the voice command using noise cancellation technology. For example, if there is strong wind noise, the reception unit can use AI to remove wind noise and recognize the voice command. For example, if background music is playing, the reception unit can use AI to filter the music and accurately recognize the voice command. In this way, by filtering out ambient noise, voice commands can be accurately recognized without being affected by noise. Some or all of the above processing in the reception unit may be performed using AI, or not using AI. For example, the reception unit can input ambient noise data into a generating AI and have the generating AI perform noise reduction.
[0067] The reception unit can estimate the user's emotions and determine the priority of voice commands based on the estimated emotions. For example, if the user is nervous, the reception unit will prioritize important commands. If the user is relaxed, the reception unit can process commands with normal priority. If the user is in a hurry, the reception unit can prioritize highly urgent commands. This ensures that important commands are prioritized by determining the priority of voice commands according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception unit may be performed using AI or not. For example, the reception unit can input the user's voice data into a generative AI and have the generative AI determine the priority of voice commands.
[0068] The reception unit can prioritize receiving commands that are highly relevant when receiving voice commands, taking into account the user's geographical location. For example, if the user is in a specific location, the reception unit can prioritize receiving commands related to that location. For example, if the user is on the move, the reception unit can prioritize receiving commands related to movement. For example, if the user is at home, the reception unit can prioritize receiving commands related to activities at home. In this way, by taking into account the user's geographical location, the reception unit can prioritize receiving commands that are highly relevant. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can input the user's geographical location data into a generating AI and have the generating AI perform the priority reception of highly relevant commands.
[0069] The reception unit can analyze the user's social media activity when receiving a voice command and suggest relevant commands. For example, if the user has posted about a specific topic on social media, the reception unit can suggest commands related to that topic. For example, if the user has participated in a specific event on social media, the reception unit can suggest commands related to that event. For example, if the user has checked in to a specific location on social media, the reception unit can suggest commands related to that location. In this way, by analyzing social media activity, commands relevant to the user can be suggested. Some or all of the above processing in the reception unit may be performed using AI, for example, or not using AI. For example, the reception unit can input the user's social media activity data into a generating AI and have the generating AI execute the suggestion of relevant commands.
[0070] The camera unit can estimate the user's emotions and adjust the framing and angle of the shot based on the estimated emotions. For example, if the user is relaxed, the camera unit can use a wide-angle, relaxed framing. If the user is excited, the camera unit can select a close-up, dynamic angle. If the user is tense, the camera unit can maintain a stable framing. This allows for the acquisition of more appropriate footage by adjusting the framing and angle of the shot according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the camera unit may be performed using AI, or not using AI. For example, the camera unit can input user emotion data into the generative AI and have the generative AI perform framing and angle adjustments.
[0071] The camera unit can predict the subject's movement and take a picture at the optimal timing. For example, the camera unit can predict the direction the subject will move and release the shutter at the optimal time. For example, the camera unit can predict the subject's speed and adjust the shutter speed to prevent blurring. For example, the camera unit can predict the subject's movement and automatically switch to continuous shooting mode. This allows the camera unit to take a picture at the optimal timing by predicting the subject's movement. Some or all of the above processes in the camera unit may be performed using AI, for example, or without AI. For example, the camera unit can input subject movement data into a generating AI and have the generating AI execute a shot at the optimal timing.
[0072] The shooting unit can automatically adjust the ambient light level and color temperature during shooting to acquire optimal images. For example, if the ambient light level is insufficient, the AI in the shooting unit can automatically increase the ISO sensitivity. For example, if the color temperature is inappropriate, the AI in the shooting unit can automatically adjust the white balance. For example, if there are extreme variations in light intensity, the AI in the shooting unit can automatically adjust the exposure. In this way, optimal images can be acquired by automatically adjusting the ambient light level and color temperature. Some or all of the above processing in the shooting unit may be performed using AI, or not. For example, the shooting unit can input ambient light level and color temperature data into a generating AI and have the generating AI perform the acquisition of optimal images.
[0073] The camera unit can estimate the user's emotions and adjust the start and stop timing of shooting based on the estimated emotions. For example, if the user is relaxed, the camera unit can adjust the start and stop timing of shooting slowly. For example, if the user is excited, the camera unit can adjust the start and stop timing of shooting quickly. For example, if the user is tense, the camera unit can stabilize the start and stop timing of shooting. By adjusting the start and stop timing of shooting according to the user's emotions, shooting can be performed at a more appropriate time. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the camera unit may be performed using AI, for example, or without AI. For example, the camera unit can input user emotion data into the generative AI and have the generative AI adjust the start and stop timing of shooting.
[0074] The shooting unit can automatically select the optimal shooting settings while considering the user's geographical location information. For example, if the user is outdoors, the AI can automatically adjust the exposure and white balance. For example, if the user is indoors, the AI can automatically adjust the ISO sensitivity and shutter speed. For example, if the user is at a specific tourist destination, the shooting unit can automatically select shooting settings appropriate for that location. In this way, the optimal shooting settings can be automatically selected by considering the user's geographical location information. Some or all of the above processing in the shooting unit may be performed using AI, for example, or without AI. For example, the shooting unit can input the user's geographical location information data into a generating AI and have the generating AI perform the automatic selection of the optimal shooting settings.
[0075] The shooting unit can analyze the user's social media activity during shooting and suggest relevant shooting scenes. For example, if the user posts about a specific topic on social media, the shooting unit can suggest shooting scenes related to that topic. For example, if the user participates in a specific event on social media, the shooting unit can suggest shooting scenes related to that event. For example, if the user checks in to a specific location on social media, the shooting unit can suggest shooting scenes related to that location. In this way, by analyzing social media activity, it is possible to suggest shooting scenes relevant to the user. Some or all of the above processing in the shooting unit may be performed using AI, for example, or not using AI. For example, the shooting unit can input the user's social media activity data into a generating AI and have the generating AI suggest relevant shooting scenes.
[0076] The editorial team can estimate the user's emotions and adjust the editing style and tempo based on the estimated emotions. For example, if the user is relaxed, the editorial team can edit at a relaxed pace. If the user is excited, the editorial team can edit in a dynamic style. If the user is tense, the editorial team can edit in a steady style. By adjusting the editing style and tempo according to the user's emotions, more appropriate editing becomes possible. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the editorial team may be performed using AI, or not using AI. For example, the editorial team can input user emotion data into a generative AI and have the generative AI adjust the editing style and tempo.
[0077] The editorial department can analyze the content of the filmed data during editing and automatically highlight important scenes. For example, in the case of a cooking video, the editorial department can highlight important cooking steps. For example, in the case of an event video, the editorial department can highlight highlight scenes. For example, in the case of an interview video, the editorial department can highlight important statements. In this way, important scenes can be automatically highlighted by analyzing the content of the filmed data. Some or all of the above processing in the editorial department may be performed using AI, for example, or without AI. For example, the editorial department can input the filmed data into a generating AI and have the generating AI perform the highlighting of important scenes.
[0078] The editorial team can automatically cut out unnecessary scenes during editing to create smoother footage. For example, in cooking videos, the editorial team can cut out waiting times and scenes of mistakes. For example, in event videos, the editorial team can cut out unnecessary scenes. For example, in interview videos, the editorial team can cut out unnecessary remarks and silences. This allows for smoother footage by automatically cutting out unnecessary scenes. Some or all of the above processing in the editorial team may be performed using AI, for example, or not. For example, the editorial team can input the shooting data into a generating AI and have the generating AI cut out unnecessary scenes.
[0079] The editorial team can estimate the user's emotions and adjust the editing order based on the estimated emotions. For example, if the user is relaxed, the editorial team can adjust the editing order in a natural flow. If the user is excited, the editorial team can edit in a dynamic order. If the user is tense, the editorial team can edit in a stable order. This allows for more appropriate editing by adjusting the editing order according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the editorial team may be performed using AI, or not using AI. For example, the editorial team can input user emotion data into a generative AI and have the generative AI adjust the editing order.
[0080] The editorial team can prioritize editing scenes that are highly relevant, taking into account the user's geographical location information. For example, the editorial team can prioritize editing scenes that the user took at a specific location. For example, the editorial team can prioritize editing scenes that the user took while traveling. For example, the editorial team can prioritize editing scenes that the user took at an event venue. This allows for the prioritization of highly relevant scenes by considering the user's geographical location information. Some or all of the above processing by the editorial team may be performed using AI, for example, or not using AI. For example, the editorial team can input the user's geographical location data into a generating AI and have the generating AI perform the priority editing of highly relevant scenes.
[0081] The editorial team can analyze users' social media activity during editing and automatically add relevant comments and effects. For example, if a user posts about a specific topic on social media, the editorial team can automatically add comments related to that topic. For example, if a user participates in a specific event on social media, the editorial team can automatically add effects related to that event. For example, if a user checks in to a specific location on social media, the editorial team can automatically add comments and effects related to that location. This allows for the automatic addition of relevant comments and effects by analyzing social media activity. Some or all of the above processing by the editorial team may be performed using AI, for example, or not. For example, the editorial team can input user social media activity data into a generating AI and have the generating AI perform the automatic addition of relevant comments and effects.
[0082] The upload unit can estimate the user's emotions and adjust the upload timing based on the estimated emotions. For example, if the user is relaxed, the upload unit can upload at an appropriate time. If the user is excited, the upload unit can upload quickly. If the user is nervous, the upload unit can upload at a steady time. By adjusting the upload timing according to the user's emotions, uploads can be performed at a more appropriate time. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the upload unit may be performed using AI, or not using AI. For example, the upload unit can input user emotion data into a generative AI and have the generative AI adjust the upload timing.
[0083] The upload unit can automatically generate video metadata and perform optimal tagging during upload. For example, the upload unit can analyze the video content and automatically tag it with relevant keywords. For example, the upload unit can automatically generate different tags for each scene in the video, making it easier for viewers to search. For example, the upload unit can automatically generate optimal tags based on the video title and description. This enables optimal tagging by automatically generating video metadata. Some or all of the above processes in the upload unit may be performed using AI, for example, or without AI. For example, the upload unit can input video content data into a generation AI and have the generation AI perform automatic metadata generation and tagging.
[0084] The upload unit can automatically adjust the video's privacy settings during upload to set the optimal visibility. For example, the upload unit can automatically select the optimal privacy settings by referring to the user's past visibility settings. For example, the upload unit can automatically adjust the visibility depending on the video's content. For example, the upload unit can analyze the user's social media activity and suggest the optimal visibility. This allows the optimal visibility to be set by automatically adjusting the video's privacy settings. Some or all of the above processes in the upload unit may be performed using AI, for example, or without AI. For example, the upload unit can input video content data into a generating AI and have the generating AI perform the automatic adjustment of privacy settings.
[0085] The upload unit can estimate the user's emotions and determine the priority of videos to upload based on the estimated emotions. For example, if the user is relaxed, the upload unit will upload videos with normal priority. If the user is excited, the upload unit can prioritize uploading important videos. If the user is stressed, the upload unit can prioritize uploading calming videos. This allows important videos to be uploaded preferentially by determining the priority of videos to upload according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the upload unit may be performed using AI or not. For example, the upload unit can input user emotion data into a generative AI and have the generative AI determine the priority of videos.
[0086] The upload unit can select the most suitable video site by considering the user's geographical location information during the upload process. For example, if the user is in a specific region, the upload unit can select a video site popular in that region. For example, if the user is traveling, the upload unit can select a video site popular in the travel destination. For example, if the user is participating in a specific event, the upload unit can select a video site related to that event. In this way, the most suitable video site can be selected by considering the user's geographical location information. Some or all of the above processing in the upload unit may be performed using AI, for example, or without AI. For example, the upload unit can input the user's geographical location data into a generating AI and have the generating AI select the most suitable video site.
[0087] The upload unit can analyze a user's social media activity during upload and automatically share it to relevant platforms. For example, if a user posts about a specific topic on social media, the upload unit can automatically share videos related to that topic. For example, if a user participates in a specific event on social media, the upload unit can automatically share videos related to that event. For example, if a user checks in to a specific location on social media, the upload unit can automatically share videos related to that location. This allows for automatic sharing to relevant platforms by analyzing social media activity. Some or all of the above processing in the upload unit may be performed using AI, for example, or without AI. For example, the upload unit can input the user's social media activity data into a generating AI and have the generating AI perform automatic sharing to relevant platforms.
[0088] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0089] The AI glasses system can also be equipped with a health management unit that monitors the user's health status. This unit can, for example, measure the user's heart rate and blood pressure and issue a warning if an abnormality is detected. It can also, for example, record the user's steps and calories burned to support daily health management. Furthermore, it can analyze the user's sleep patterns and provide advice to promote high-quality sleep. This allows the AI glasses system to monitor the user's health status in real time and support health management.
[0090] The AI glasses system can also include a reminder unit that estimates the user's emotions and provides appropriate reminders based on those emotions. For example, if the user is feeling stressed, the reminder unit can suggest taking a break to relax. If the user is concentrating, the reminder unit can provide reminders for important tasks. If the user is tired, the reminder unit can provide reminders to go to bed early. By providing appropriate reminders according to the user's emotions, the quality of daily life can be improved.
[0091] AI glasses systems can also include a recommendation system that suggests content based on the user's hobbies and interests. For example, if the user likes music, the recommendation system can provide the latest music trends and recommended playlists. If the user likes movies, the recommendation system can provide the latest movie information and recommended movies. If the user likes sports, the recommendation system can provide the latest match results and recommended sporting events. This can enhance the user's entertainment experience by recommending content based on their hobbies and interests.
[0092] The AI glasses system can also include a feedback unit that estimates the user's emotions and provides appropriate feedback based on those emotions. For example, if the user is feeling anxious, the feedback unit can provide an encouraging message. If the user is happy, the feedback unit can provide a congratulatory message. If the user is feeling down, the feedback unit can provide uplifting advice. In this way, by providing appropriate feedback according to the user's emotions, it can provide emotional support to the user.
[0093] The AI glasses system can also be equipped with a learning support unit to assist the user's learning. For example, if the user is learning a new language, the learning support unit can provide advice on pronunciation and grammar. If the user is studying for an exam, the learning support unit can provide efficient learning methods and key points. If the user wants to improve a hobby skill, the learning support unit can provide relevant learning materials and practice methods. This supports the user's learning, thereby promoting skill improvement and knowledge acquisition.
[0094] The AI glasses system can also include an exercise unit that estimates the user's emotions and suggests appropriate exercises based on those emotions. For example, if the user is feeling stressed, the exercise unit might suggest yoga or stretching to relax. If the user is feeling energetic, it might suggest running or high-intensity training. If the user is feeling tired, it might suggest light walking or relaxation exercises. This allows the system to support a healthy lifestyle by suggesting appropriate exercises according to the user's emotions.
[0095] The AI glasses system can also be equipped with a schedule management unit to support the user's schedule management. For example, the schedule management unit can automatically organize the user's schedule and remind them of important events and tasks. For example, the schedule management unit can suggest the optimal time allocation based on the user's schedule. For example, the schedule management unit can provide traffic and weather information according to the user's schedule. This allows for efficient time management by supporting the user's schedule management.
[0096] The AI glasses system can also include a music recommendation unit that estimates the user's emotions and recommends appropriate music based on those emotions. For example, if the user is relaxed, the music recommendation unit can recommend relaxing music. If the user is excited, the music recommendation unit can recommend energetic music. If the user is sad, the music recommendation unit can recommend mood-calming music. This can improve the music experience by recommending appropriate music according to the user's emotions.
[0097] The AI glasses system can also be equipped with a travel support unit to assist users during their travels. This unit could, for example, recommend tourist attractions and restaurants based on the user's current location. It could also suggest optimal routes based on the user's travel schedule. Furthermore, it could provide customized travel plans tailored to the user's preferences. This would allow the system to support users' travels and provide a more fulfilling travel experience.
[0098] The AI glasses system can also include a relaxation unit that estimates the user's emotions and suggests appropriate relaxation methods based on those emotions. For example, if the user is feeling stressed, the relaxation unit can suggest meditation or deep breathing techniques. If the user is feeling tired, the relaxation unit can suggest relaxing music or aromatherapy. If the user is feeling anxious, the relaxation unit can suggest relaxing massage or a warm bath. In this way, by suggesting appropriate relaxation methods according to the user's emotions, it can support mental and physical refreshment.
[0099] The following briefly describes the processing flow for example form 2.
[0100] Step 1: The reception desk receives voice commands. Voice commands include phrases such as "Start shooting" or "Take a picture of the food." The reception desk uses voice recognition technology to analyze the voice commands and generate appropriate instructions. Step 2: The shooting unit takes pictures based on the voice commands received by the reception unit. The shooting unit has a built-in camera and captures images from the user's point of view. It can capture stable, wide-angle images using image stabilization and a wide-angle lens. Step 3: The editorial team analyzes the footage shot by the camera crew and performs optimal editing. Using AI, they analyze the footage to highlight important scenes and cut out unnecessary parts. Furthermore, the AI automatically adds commentary, music, and sound effects to create a video that is easy to understand and engaging for viewers. Step 4: The uploading unit automatically uploads the videos edited by the editorial team to the video site. It accesses the video site via an internet connection and uploads the videos. Furthermore, it automatically generates video titles and descriptions to provide engaging content for viewers.
[0101] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0102] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0103] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0104] Each of the multiple elements described above, including the reception unit, shooting unit, editing unit, and upload unit, is implemented by, for example, at least one of the smart device 14 and the data processing unit 12. For example, the reception unit receives voice commands via the microphone 38B and control unit 46A of the smart device 14. The shooting unit captures video from the user's point of view using the camera 42 of the smart device 14. The editing unit analyzes the captured data using the specific processing unit 290 of the data processing unit 12 and performs optimal editing. The upload unit connects to the internet via the communication I / F 26 of the data processing unit 12 and uploads the edited video to a video site. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0105] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0106] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0107] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0108] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0109] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0110] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0111] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0112] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0113] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0114] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0115] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0116] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0117] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0118] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0119] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0120] Each of the multiple elements described above, including the reception unit, shooting unit, editing unit, and upload unit, is implemented by, for example, at least one of the smart glasses 214 and the data processing unit 12. For example, the reception unit receives voice commands via the microphone 238 and control unit 46A of the smart glasses 214. The shooting unit captures video from the user's point of view using the camera 42 of the smart glasses 214. The editing unit analyzes the captured data using the specific processing unit 290 of the data processing unit 12 and performs optimal editing. The upload unit connects to the internet via the communication I / F 26 of the data processing unit 12 and uploads the edited video to a video site. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0121] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0122] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0123] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0124] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0125] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0126] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0127] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0128] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0129] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0130] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0131] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0132] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0133] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0134] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0135] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0136] Each of the multiple elements described above, including the reception unit, shooting unit, editing unit, and upload unit, is implemented by, for example, at least one of the headset terminal 314 and the data processing unit 12. For example, the reception unit receives voice commands via the microphone 238 and control unit 46A of the headset terminal 314. The shooting unit captures video from the user's point of view using the camera 42 of the headset terminal 314. The editing unit analyzes the captured data using the specific processing unit 290 of the data processing unit 12 and performs optimal editing. The upload unit connects to the internet via the communication I / F 26 of the data processing unit 12 and uploads the edited video to a video site. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0137] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0138] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0139] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0140] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0141] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0142] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0143] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0144] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0145] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0146] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0147] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0148] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0149] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0150] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0151] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0152] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0153] Each of the multiple elements described above, including the reception unit, shooting unit, editing unit, and upload unit, is implemented by, for example, at least one of the robot 414 and the data processing unit 12. For example, the reception unit receives voice commands via the microphone 238 and control unit 46A of the robot 414. The shooting unit captures video from the user's point of view using the camera 42 of the robot 414. The editing unit analyzes the captured data using the specific processing unit 290 of the data processing unit 12 and performs optimal editing. The upload unit connects to the internet via the communication I / F 26 of the data processing unit 12 and uploads the edited video to a video site. The correspondence between each unit and the devices and control units is not limited to the example described above and can be modified in various ways.
[0154] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0155] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0156] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0157] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0158] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0159] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0160] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0161] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0162] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0163] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0164] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0165] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0166] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0167] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0168] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0169] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0170] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0171] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0172] (Note 1) A reception area that accepts voice commands, A shooting unit that takes pictures based on voice commands received by the aforementioned reception unit, The editing department analyzes the data captured by the aforementioned shooting department and performs optimal editing. The system includes an upload unit that automatically uploads videos edited by the aforementioned editorial department to a video site. A system characterized by the following features. (Note 2) The aforementioned reception unit is It estimates the user's emotions and adjusts the accuracy of voice command recognition based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned reception unit is When a voice command is received, the system refers to the user's past command history to perform the optimal command interpretation. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned reception unit is When receiving voice commands, ambient noise is filtered and noise is removed. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned reception unit is It estimates the user's emotions and prioritizes voice commands based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned reception unit is When receiving voice commands, the system prioritizes accepting commands that are highly relevant, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned reception unit is When a voice command is received, the system analyzes the user's social media activity and suggests relevant commands. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned imaging unit is It estimates the user's emotions and adjusts the framing and angle of the shot based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned imaging unit is During shooting, predict the subject's movement and take the picture at the optimal timing. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned imaging unit is During shooting, the camera automatically adjusts the ambient light level and color temperature to capture the optimal image. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned imaging unit is It estimates the user's emotions and adjusts the start and stop timing of recording based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned imaging unit is During shooting, the system automatically selects the optimal shooting settings, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned imaging unit is During shooting, the system analyzes the user's social media activity and suggests relevant shooting scenes. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned editorial department, It estimates the user's emotions and adjusts the editing style and tempo based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned editorial department, During editing, the system analyzes the content of the footage and automatically highlights important scenes. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned editorial department, During editing, unnecessary scenes are automatically cut out, resulting in a smoother video. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned editorial department, It estimates the user's emotions and adjusts the editing order based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned editorial department, During editing, the system prioritizes editing scenes that are highly relevant, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned editorial department, During editing, the system analyzes the user's social media activity and automatically adds relevant comments and effects. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned upload unit, It estimates the user's emotions and adjusts the upload timing based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned upload unit, During upload, video metadata is automatically generated and optimal tags are applied. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned upload unit, When uploading, the video's privacy settings are automatically adjusted to set the optimal level of visibility. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned upload unit, It estimates user sentiment and prioritizes uploading videos based on the estimated user sentiment. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned upload unit, When uploading, the system selects the most suitable video site by considering the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned upload unit, When uploading, the system analyzes the user's social media activity and automatically shares it to relevant platforms. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]
[0173] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. A reception area that accepts voice commands, A shooting unit that takes pictures based on voice commands received by the aforementioned reception unit, The editing department analyzes the data captured by the aforementioned shooting department and performs optimal editing. The system includes an upload unit that automatically uploads videos edited by the aforementioned editorial department to a video site. A system characterized by the following features.
2. The aforementioned reception unit is It estimates the user's emotions and adjusts the accuracy of voice command recognition based on the estimated emotions. The system according to feature 1.
3. The aforementioned reception unit is When a voice command is received, the system refers to the user's past command history to perform the optimal command interpretation. The system according to feature 1.
4. The aforementioned reception unit is When receiving voice commands, ambient noise is filtered and noise is removed. The system according to feature 1.
5. The aforementioned reception unit is It estimates the user's emotions and prioritizes voice commands based on those emotions. The system according to feature 1.
6. The aforementioned reception unit is When receiving voice commands, the system prioritizes accepting commands that are highly relevant, taking into account the user's geographical location. The system according to feature 1.
7. The aforementioned reception unit is When a voice command is received, the system analyzes the user's social media activity and suggests relevant commands. The system according to feature 1.
8. The aforementioned imaging unit is It estimates the user's emotions and adjusts the framing and angle of the shot based on the estimated emotions. The system according to feature 1.
9. The aforementioned imaging unit is During shooting, predict the subject's movement and take the picture at the optimal timing. The system according to feature 1.
10. The aforementioned imaging unit is During shooting, the camera automatically adjusts the ambient light level and color temperature to capture the optimal image. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A