system
The system addresses the issue of non-preferred commentator settings by enabling personalized voice adjustments, focused commentary, and interactive communication, enhancing user engagement and enjoyment in sports viewing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
In sports viewing, the voice of the commentator and the content of the live broadcast cannot be adjusted according to the user's preferences, limiting the quality of the viewing experience.
A system comprising a reception unit, adjustment unit, focus unit, communication unit, and provision unit that allows users to adjust the quality, volume, color, and intensity of the commentator's voice, provides commentary focused on specific players, enables real-time communication with commentators, and generates virtual spectators for a personalized viewing experience.
The system provides tailored commentary and live broadcasts that enhance user engagement and enjoyment by allowing personalized adjustments and interactive features, preventing viewer dropout and attracting new audiences.
Smart Images

Figure 2026072369000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there is a problem that in sports viewing, the voice of the commentator and the content of the live broadcast cannot be adjusted according to the user's preferences, and the quality of the viewing experience is limited.
[0005] The system according to the embodiment aims to provide explanations and live broadcasts according to the user's preferences in sports viewing.
Means for Solving the Problems
[0006] The system according to this embodiment comprises a reception unit, an adjustment unit, a focus unit, a communication unit, and a provision unit. The reception unit receives user input. The adjustment unit adjusts the quality, volume, color, and intensity of the commentator's voice based on the information received by the reception unit. The focus unit provides commentary focused on a specific player based on the information adjusted by the adjustment unit. The communication unit communicates with the commentator based on the commentary performed by the focus unit. The provision unit provides the results of the communication performed by the communication unit. [Effects of the Invention]
[0007] The system according to this embodiment can provide commentary and live broadcasts tailored to the user's preferences when watching sports. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9]This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The sports viewing personalization system according to an embodiment of the present invention is a system that personalizes sports viewing and provides a new entertainment experience. This system allows users to adjust the quality, volume, color, and intensity of the commentator's voice while watching sports. For example, the volume of the commentator can be reduced, and the audio can be set to focus on the atmosphere of the venue. This allows users to enjoy watching sports in an audio environment tailored to their preferences. Next, users can request commentary that focuses on their favorite players. For example, by having the commentary focus on scenes where a specific player is playing, users can enjoy that player's performance more deeply. Furthermore, users can use a communication function with the commentator to converse with the commentator in real time. For example, users can ask questions to the commentator or receive comments from the commentator. This allows users to enjoy two-way communication rather than just passively listening to commentary. In addition, even when watching alone, users can watch while engaging in conversations as if they were watching with multiple people. For example, by watching while conversing with multiple virtual spectators generated by AI, users can have an experience as if they were watching with friends. This system can prevent viewers from dropping out due to commentators, attract new audiences, and increase the overall value of sports. For example, users who find commentators' voices too loud can adjust the volume for a more comfortable viewing experience. Furthermore, providing commentary focused on specific players can attract fans of those players. Additionally, offering a group viewing experience allows users who feel lonely watching alone to enjoy themselves. In this way, the present invention realizes a new entertainment experience for sports viewing by providing a personalized sports viewing experience tailored to user needs. Thus, the sports viewing personalization system can provide a personalized sports viewing experience tailored to user needs.
[0029] The sports viewing personalization system according to this embodiment comprises a reception unit, an adjustment unit, a focus unit, a communication unit, and a provision unit. The reception unit receives user input. User input includes, but is not limited to, voice input, text input, and gesture input. The reception unit analyzes the user's voice input using, for example, speech recognition technology and converts it into text data. The reception unit can also directly receive text input. Furthermore, the reception unit can analyze the user's gesture input using gesture recognition technology and execute a corresponding command. For example, if the user instructs the reception unit by voice to "lower the volume of the commentary," the reception unit analyzes the voice and generates a command to lower the volume of the commentary. The adjustment unit adjusts the quality, quantity, color, and intensity of the commentator's voice based on the information received by the reception unit. The adjustment unit adjusts the tone and volume of the commentator's voice using, for example, voice filtering technology. The adjustment unit can also add echo and reverb to the commentator's voice using effect technology. Furthermore, the adjustment unit can adjust the emotion of the commentator's voice using emotion expression technology. For example, the adjustment unit adjusts the commentator's voice tone to a calmer one, allowing users to relax while watching the game. The focus unit provides commentary focused on specific players based on the information adjusted by the adjustment unit. For example, the focus unit analyzes player performance data and provides commentary centered on scenes where a particular player is playing. The focus unit can also provide commentary focused on the user's favorite players based on user preference information. Furthermore, the focus unit can detect important moments in the match and provide commentary focused on those moments. For example, the focus unit can focus on the moment a particular player scores a goal and provide a detailed explanation of that play. The communication unit communicates with the commentator based on the commentary provided by the focus unit. For example, the communication unit uses voice dialogue technology to enable real-time conversations between the user and the commentator. The communication unit can also use text chat technology to enable text-based communication between the user and the commentator.Furthermore, the communication unit can also use gesture recognition technology to convey the user's gestures to the commentator. For example, the communication unit can have the user ask a question to the commentator and provide the commentator's answer to that question in real time. The delivery unit provides the results of the communication conducted by the communication unit. The delivery unit provides the user with, for example, adjusted audio or generated virtual spectator conversations. The delivery unit can also deliver information to the user's device in real time. Furthermore, the delivery unit can provide customized information tailored to the user's preferences. For example, the delivery unit can provide information about a specific player requested by the user. As a result, the sports viewing personalization system according to the embodiment can adjust the quality, volume, color, and intensity of the commentator's voice based on the user's input, provide commentary focused on a specific player, and communicate with the commentator.
[0030] The reception unit receives user input. User input includes, but is not limited to, voice input, text input, and gesture input. For example, the reception unit analyzes the user's voice input using speech recognition technology and converts it into text data. Specifically, speech recognition technology analyzes the user's speech in real time and converts the speech waveform into a digital signal. Then, it converts the speech data into text data using an acoustic model and a language model. The acoustic model extracts the features of the speech, and the language model selects the most appropriate words and phrases based on the context. This ensures that the user's voice input is accurately converted into text data. The reception unit can also directly accept text input. The user enters text using a keyboard or touchscreen, and that text data is sent to the system. Furthermore, the reception unit can analyze the user's gesture input using gesture recognition technology and execute corresponding commands. Gesture recognition technology detects the user's movements using cameras and sensors and analyzes the movement patterns. For example, if the user waves their hand, that movement is set to correspond to a specific command. This allows the user to operate the system intuitively. For example, if a user gives a voice command saying, "Lower the volume of the commentary," the reception desk analyzes the voice and generates a command to lower the volume. This allows users to easily operate the system and obtain a viewing experience tailored to their preferences.
[0031] The adjustment unit adjusts the quality, volume, tone, and intensity of the commentator's voice based on the information received by the reception unit. For example, the adjustment unit can adjust the tone and volume of the commentator's voice using voice filtering technology. Voice filtering technology improves the quality of the sound by emphasizing or attenuating specific frequency bands. For example, emphasizing low frequency bands can give the commentator's voice more depth. The adjustment unit can also add echo and reverb to the commentator's voice using effects technology. Echo and reverb give the sound a sense of spatial breadth and enhance the feeling of presence. Furthermore, the adjustment unit can adjust the emotion in the commentator's voice using emotion expression technology. Emotion expression technology adds emotion to the commentator's voice by adjusting the pitch, tempo, and volume of the voice. For example, the tone of the commentator's voice can be changed to a calmer tone so that the user can watch the game in a relaxed manner. In this way, the adjustment unit can customize the commentator's voice to suit the user's preferences and provide a more personalized viewing experience.
[0032] The focus unit provides commentary focused on specific players based on information adjusted by the adjustment unit. For example, the focus unit analyzes player performance data and provides commentary centered on the moments when a particular player is playing. Performance data includes player distance covered, number of sprints, and shooting success rate. This data is analyzed in real time to evaluate player performance. For example, it can focus on the moment a particular player scores a goal and provide a detailed explanation of that play. The focus unit can also provide commentary focused on the user's favorite players based on user preference information. User preference information is collected based on past viewing history and user input. This allows users to watch the game focusing on the play of their favorite players. Furthermore, the focus unit can detect important moments in the match and provide commentary focused on those moments. Important moments are detected based on events such as goals, assists, and fouls. This allows users to enjoy important moments without missing the highlights of the match.
[0033] The Communications Department communicates with commentators based on the live commentary provided by the Focus Department. The Communications Department, for example, uses voice dialogue technology to enable real-time conversations between users and commentators. This voice dialogue technology analyzes the user's voice input and generates appropriate responses. For example, if a user asks, "What are this player's past statistics?", the Communications Department analyzes the question, and the commentator provides information about the player's past statistics. The Communications Department can also use text chat technology for text-based communication between users and commentators. This technology analyzes the text entered by the user and generates appropriate responses. Furthermore, the Communications Department can use gesture recognition technology to convey the user's gestures to the commentators. For example, if a user raises their hand, this gesture is interpreted as indicating the intent of a question, and the commentator provides an answer. This allows users to enjoy interactive communication with commentators.
[0034] The service provider provides the results of communication conducted by the communication provider. For example, the service provider provides users with adjusted audio and conversations with generated virtual spectators. Virtual spectators are generated using AI and can enjoy watching the game together with the user. Virtual spectators provide appropriate comments and cheers in response to the user's reactions. This allows users to have an experience as if they were watching with friends, even when watching alone. The service provider can also deliver information to the user's device in real time. For example, it can deliver the latest match information and commentary in real time to smartphones and tablets. Furthermore, the service provider can provide customized information tailored to the user's preferences. For example, it can provide information about a specific player that the user has requested. This allows users to obtain information that matches their interests and enjoy a more fulfilling viewing experience.
[0035] The system includes a virtual spectator unit that generates conversations with virtual spectators. The virtual spectator unit can generate virtual spectators using, for example, AI characters. The virtual spectator unit can also generate the actions and conversations of virtual spectators using simulation models. For example, the virtual spectator unit can generate scenarios where an AI character converses with the user while watching the game. Furthermore, the virtual spectator unit can use simulation models to provide an experience as if multiple virtual spectators were watching the game with the user. For example, the virtual spectator unit can have an AI character generate conversations in real time according to the progress of the match. This allows the user to experience watching the game with multiple people by generating conversations with virtual spectators.
[0036] The adjustment unit can reduce the volume of commentary and set the audio to primarily reflect the atmosphere of the venue. For example, the adjustment unit can reduce the volume of commentary and emphasize crowd cheers and ambient sounds. The adjustment unit can also use audio filtering technology to realistically reproduce the atmosphere of the venue. For example, by emphasizing crowd cheers and reducing the volume of commentary, the adjustment unit can provide users with a sense of presence as if they were actually at the venue. In this way, by reducing the volume of commentary and setting the audio to primarily reflect the atmosphere of the venue, users can enjoy watching sports in an audio environment tailored to their preferences.
[0037] The focus unit can provide commentary centered on moments in which a specific player is playing. For example, the focus unit can analyze player performance data to detect moments in which a particular player is playing. Furthermore, based on user preference information, the focus unit can provide commentary focused on the user's favorite players. For instance, the focus unit can focus on the moment a specific player scores a goal and provide a detailed commentary on that play. This allows users to enjoy that player's performance more deeply by focusing the commentary on moments in which they are playing.
[0038] The communication section allows users to ask questions to commentators and receive comments from them. The communication section enables real-time conversations between users and commentators using, for example, voice dialogue technology. It can also facilitate text-based communication between users and commentators using text chat technology. For instance, the communication section allows users to ask questions to commentators and provides real-time responses from the commentators. This allows users to enjoy two-way communication by asking questions and receiving comments from commentators.
[0039] The service provider can offer users pre-tuned audio and generated virtual spectator conversations. For example, the service provider can deliver pre-tuned audio to the user's device in real time. The service provider can also offer users generated virtual spectator conversations. For example, the service provider can provide information about a specific player requested by the user. By providing users with pre-tuned audio and generated virtual spectator conversations, users can enjoy watching sports tailored to their preferences.
[0040] The reception desk can analyze a user's past input history and select the optimal input method. For example, the reception desk can prioritize suggesting input methods (voice, text, etc.) that the user has frequently used in the past. It can also predict and suggest input methods to be used during specific time periods based on the user's past input history. Furthermore, the reception desk can provide an auto-completion function to reduce input effort by referencing the user's past input. For example, if the reception desk has frequently used voice input in the past, it will prioritize suggesting voice input. Similarly, if the reception desk has used text input during specific time periods in the past, it can suggest text input during those times. Finally, the reception desk can provide an auto-completion function based on the user's past input to reduce input effort. In this way, by analyzing the user's past input history, the optimal input method can be selected, improving user convenience.
[0041] The reception system can filter input based on the user's current viewing status and areas of interest. For example, if a user is watching a particular sport, the reception system will prioritize receiving only input related to that sport. Similarly, if a user is interested in a particular player, the reception system can prioritize receiving information related to that player. Furthermore, if a user is watching a particular match, the reception system can prioritize receiving only input related to that match. For example, if a user is watching a soccer match, the reception system will prioritize receiving input related to soccer. Similarly, if a user is interested in a particular player, the reception system can prioritize receiving information related to that player. Furthermore, if a user is watching a particular match, the reception system can prioritize receiving input related to that match. This allows the system to prioritize receiving highly relevant information by filtering based on the user's current viewing status and areas of interest.
[0042] The reception desk can prioritize receiving highly relevant input by considering the user's geographical location. For example, if the user is in a specific region, the reception desk will prioritize receiving information related to that region. Similarly, if the user is in a specific stadium, the reception desk can prioritize receiving information related to that stadium. Furthermore, if the user is in a specific city, the reception desk can prioritize receiving information related to that city. This allows the reception desk to prioritize receiving highly relevant information by considering the user's geographical location.
[0043] The reception desk can analyze the user's social media activity when receiving input and accept relevant input. For example, if the reception desk mentions a specific athlete on social media, it will prioritize receiving information related to that athlete. Similarly, if the reception desk mentions a specific match on social media, it can prioritize receiving information related to that match. Furthermore, if the reception desk mentions a specific sport on social media, it can prioritize receiving information related to that sport. This allows the reception desk to prioritize receiving relevant information by analyzing the user's social media activity.
[0044] The adjustment unit can select the optimal adjustment method by referring to the user's past viewing history during the adjustment process. For example, the adjustment unit can select the optimal adjustment method based on the commentary style the user has preferred in the past. Furthermore, the adjustment unit can prioritize the style of a specific commentator based on the user's past viewing history. In addition, the adjustment unit can analyze the user's past viewing history and select the most appropriate quality and quantity of commentary. This allows the system to select the optimal adjustment method by referring to the user's past viewing history, thereby improving user convenience.
[0045] The adjustment unit can customize the adjustment methods based on the user's current viewing situation. For example, if the user is watching a specific match, the adjustment unit can customize the quality and quantity of commentary to be optimal for that match. It can also customize commentary related to a specific player if the user is focusing on that player. Furthermore, if the user is watching a specific sport, the adjustment unit can customize the quality and quantity of commentary to be optimal for that sport. This allows the system to provide optimal commentary by customizing the adjustment methods based on the user's current viewing situation.
[0046] The adjustment unit can select the optimal adjustment method during adjustment, taking into account the user's geographical location information. For example, if the user is in a specific region, the adjustment unit can prioritize providing explanations related to that region. It can also prioritize providing explanations related to a specific stadium if the user is in that stadium. Furthermore, if the user is in a specific city, the adjustment unit can prioritize providing explanations related to that city. This allows the system to select the optimal adjustment method and improve user convenience by considering the user's geographical location information.
[0047] The adjustment unit can analyze the user's social media activity during the adjustment process and propose adjustment methods. For example, if the user mentions a specific player on social media, the adjustment unit can prioritize providing commentary related to that player. Similarly, if the user mentions a specific match on social media, the adjustment unit can prioritize providing commentary related to that match. Furthermore, if the user mentions a specific sport on social media, the adjustment unit can prioritize providing commentary related to that sport. This allows the system to analyze the user's social media activity, propose the most suitable adjustment method, and improve user convenience.
[0048] The focus unit can select the optimal focus method by referring to the user's past viewing history when focusing. For example, the focus unit can select the optimal focus method based on the player or scene the user has liked in the past. The focus unit can also prioritize focusing on specific players or scenes based on the user's past viewing history. Furthermore, the focus unit can analyze the user's past viewing history and focus on the most suitable player or scene. For example, the focus unit can select the optimal focus method based on the player or scene the user has liked in the past. Furthermore, the focus unit can prioritize focusing on specific players or scenes based on the user's past viewing history. Furthermore, the focus unit can analyze the user's past viewing history and focus on the most suitable player or scene. By referring to the user's past viewing history, the optimal focus method can be selected, improving user convenience.
[0049] The focus unit can customize the means of focusing based on the user's current viewing situation. For example, if the user is watching a specific match, the focus unit will focus on the most suitable players or scenes for that match. It can also focus on scenes related to a specific player if the user is focusing on that player. Furthermore, if the user is watching a specific sport, the focus unit can focus on the most suitable players or scenes for that sport. This allows the system to provide optimal focus by customizing the means of focusing based on the user's current viewing situation.
[0050] The focus unit can select the optimal focus method when focusing, taking into account the user's geographical location. For example, if the user is in a specific region, the focus unit will prioritize focusing on players or scenes related to that region. Similarly, if the user is in a specific stadium, the focus unit can prioritize focusing on players or scenes related to that stadium. Furthermore, if the user is in a specific city, the focus unit can prioritize focusing on players or scenes related to that city. This allows the system to select the optimal focus method by considering the user's geographical location, thereby improving user convenience.
[0051] The focus unit can analyze the user's social media activity and suggest a method of focus when focusing. For example, if the user mentions a specific athlete on social media, the focus unit will prioritize focusing on scenes related to that athlete. Similarly, if the user mentions a specific match on social media, the focus unit can prioritize focusing on scenes related to that match. Furthermore, if the user mentions a specific sport on social media, the focus unit can prioritize focusing on scenes related to that sport. This allows the system to analyze the user's social media activity, suggest the optimal method of focus, and improve user convenience.
[0052] The communication department can select the optimal method of communication by referring to the user's past communication history. For example, the communication department can select the optimal method based on the communication style the user has preferred in the past. Furthermore, the communication department can prioritize specific tones and styles based on the user's past communication history. In addition, the communication department can analyze the user's past communication history and select the most appropriate communication method. This allows the system to select the optimal communication method by referring to the user's past communication history, thereby improving user convenience.
[0053] The communications department can customize the means of communication based on the user's current viewing situation. For example, if the user is watching a particular match, the communications department can customize the communication method best suited to that match. Furthermore, if the user is focusing on a particular player, the communications department can customize communication related to that player. Additionally, if the user is watching a particular sport, the communications department can customize communication methods best suited to that sport. This allows for optimal communication by customizing the means of communication based on the user's current viewing situation.
[0054] The communications department can select the optimal method of communication by considering the user's geographical location. For example, if the user is in a specific region, the communications department can prioritize communications related to that region. Similarly, if the user is in a specific stadium, the communications department can prioritize communications related to that stadium. Furthermore, if the user is in a specific city, the communications department can prioritize communications related to that city. This allows the communications department to select the optimal communication method by considering the user's geographical location, thereby improving user convenience.
[0055] The communications department can analyze users' social media activity and suggest appropriate communication methods during communication. For example, if a user mentions a specific athlete on social media, the communications department will prioritize communications related to that athlete. Similarly, if a user mentions a specific match on social media, the communications department can prioritize communications related to that match. Furthermore, if a user mentions a specific sport on social media, the communications department can prioritize communications related to that sport. By analyzing users' social media activity, the communications department can suggest optimal communication methods and improve user convenience.
[0056] The service provider can select the optimal delivery method by referring to the user's past viewing history at the time of delivery. For example, the service provider can select the optimal delivery method based on the information format the user has preferred in the past. Furthermore, the service provider can prioritize the selection of a specific information format based on the user's past viewing history. In addition, the service provider can analyze the user's past viewing history and select the most suitable information format. This allows the service provider to select the optimal delivery method by referring to the user's past viewing history, thereby improving user convenience.
[0057] The service provider can customize the means of delivery based on the user's current viewing situation. For example, if the user is watching a particular match, the service provider can customize the information format to be best suited to that match. Furthermore, if the user is focusing on a particular player, the service provider can customize the information format to be best suited to that player. In addition, if the user is watching a particular sport, the service provider can customize the information format to be best suited to that sport. This allows the service provider to provide optimal information by customizing the means of delivery based on the user's current viewing situation.
[0058] The service provider can select the optimal delivery method by considering the user's geographical location at the time of delivery. For example, if the user is in a specific region, the service provider can prioritize providing information related to that region. Similarly, if the user is in a specific stadium, the service provider can prioritize providing information related to that stadium. Furthermore, if the user is in a specific city, the service provider can prioritize providing information related to that city. This allows the service provider to select the optimal delivery method by considering the user's geographical location, thereby improving user convenience.
[0059] The service provider can analyze the user's social media activity at the time of delivery and propose a delivery method. For example, if the service provider mentions a specific athlete on social media, it can prioritize providing information related to that athlete. Similarly, if the service provider mentions a specific match on social media, it can prioritize providing information related to that match. Furthermore, if the service provider mentions a specific sport on social media, it can prioritize providing information related to that sport. By analyzing the user's social media activity, the service provider can propose the optimal delivery method and improve user convenience.
[0060] The virtual spectator unit can select the most suitable conversation content by referring to the user's past viewing history when generating conversations for virtual spectators. For example, the virtual spectator unit can select the most suitable conversation content based on the conversation content the user has enjoyed in the past. Furthermore, the virtual spectator unit can prioritize the selection of specific conversation content from the user's past viewing history. In addition, the virtual spectator unit can analyze the user's past viewing history and select the most suitable conversation content. This allows for the selection of the most suitable conversation content by referring to the user's past viewing history, thereby improving user convenience.
[0061] The virtual spectator unit can customize conversation content based on the user's current viewing situation when generating conversations with virtual spectators. For example, if the user is watching a specific match, the virtual spectator unit will customize conversation content related to that match. It can also customize conversation content related to a specific player if the user is focusing on that player. Furthermore, if the virtual spectator unit is watching a specific sport, it can customize conversation content related to that sport. This allows the system to provide optimal conversations by customizing conversation content based on the user's current viewing situation.
[0062] The virtual spectator system can select the most relevant conversation content when generating conversations for virtual spectators, taking into account the user's geographical location. For example, if the user is in a specific region, the system will prioritize providing conversation content related to that region. It can also prioritize providing conversation content related to a specific stadium if the user is in that stadium. Furthermore, if the user is in a specific city, the system can prioritize providing conversation content related to that city. This allows the system to select the most relevant conversation content by considering the user's geographical location, thereby improving user convenience.
[0063] The virtual spectator system can analyze a user's social media activity to select the most relevant conversation content when generating conversations for virtual spectators. For example, if a user mentions a specific player on social media, the system will prioritize providing conversation content related to that player. Similarly, if a user mentions a specific match on social media, the system can prioritize providing conversation content related to that match. Furthermore, if a user mentions a specific sport on social media, the system can prioritize providing conversation content related to that sport. This allows the system to analyze a user's social media activity, select the most relevant conversation content, and improve user experience.
[0064] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0065] A personalized sports viewing system can provide customized content based on the user's hobbies and interests. For example, if a user is interested in movies or music in addition to a particular sport, the system can provide information and content related to those interests while they are watching a sport. Similarly, if a user is interested in a specific region or culture, the system can provide information related to that region or culture. Furthermore, if a user is interested in a specific event or festival, the system can provide information related to that event. This allows for a customized viewing experience based on the user's hobbies and interests.
[0066] A personalized sports viewing system can analyze a user's viewing history and provide information related to their favorite matches and players in the past. For example, it can provide highlights of matches the user has watched in the past or compilations of specific players' plays. It can also suggest the most suitable commentator based on the user's preferred commentary style and announcers. Furthermore, it can provide information related to events and festivals the user has attended in the past. This allows for a customized viewing experience based on the user's viewing history.
[0067] A personalized sports viewing system can customize the information provided during a game by taking into account the user's geographical location. For example, if the user is in a specific region, it can provide sports news and event information related to that region. If the user is at a specific stadium, it can provide information and history related to that stadium. Furthermore, if the user is in a specific city, it can provide tourist and cultural information related to that city. This allows for a customized viewing experience based on the user's geographical location.
[0068] A personalized sports viewing system can analyze a user's social media activity and customize the information provided during viewing. For example, if a user mentions a specific player on social media, it can prioritize providing information related to that player. Similarly, if a user mentions a specific match on social media, it can prioritize providing information related to that match. Furthermore, if a user mentions a specific sport on social media, it can prioritize providing information related to that sport. This allows for a customized viewing experience based on the user's social media activity.
[0069] A personalized sports viewing system can monitor a user's behavior while watching a game and adjust the viewing experience accordingly. For example, if a user frequently gets up from their seat, it can display alerts to ensure they don't miss important moments. It can also display alerts encouraging users to take breaks if they are sitting in the same position for extended periods. Furthermore, if a user repeatedly engages in certain behaviors while watching, the system can provide information and content related to those behaviors. This allows for a customized viewing experience based on the user's actions during the game.
[0070] The following briefly describes the processing flow for example form 1.
[0071] Step 1: The reception unit receives user input. User input includes voice input, text input, and gesture input. For example, voice recognition technology can be used to analyze the user's voice input and convert it into text data. It can also directly accept text input, and gesture recognition technology can be used to analyze the user's gesture input and execute the corresponding command. Step 2: The adjustment unit adjusts the quality, volume, color, and intensity of the commentator's voice based on the information received by the reception unit. For example, it can use voice filtering technology to adjust the tone and volume of the commentator's voice, and effect technology to add echo and reverb to the commentator's voice. Furthermore, it can also adjust the emotion of the commentator's voice using emotion expression technology. Step 3: The focus unit provides commentary focused on a specific player based on the information adjusted by the adjustment unit. For example, it can analyze player performance data and provide commentary focusing on scenes where a particular player is playing. It can also provide commentary focused on the user's favorite player based on user preference information. Furthermore, it can detect important moments in the match and provide commentary focused on those moments. Step 4: The communication unit communicates with the commentator based on the live commentary provided by the focus unit. For example, voice dialogue technology can be used to enable real-time conversations between the user and the commentator, and text chat technology can be used for text-based communication between the user and the commentator. Furthermore, gesture recognition technology can be used to convey the user's gestures to the commentator. Step 5: The delivery unit provides the results of the communication conducted by the communication unit. For example, it can provide users with coordinated audio or generated virtual spectator conversations, and can also deliver information to the user's device in real time. Furthermore, it can provide customized information tailored to the user's preferences.
[0072] (Example of form 2) The sports viewing personalization system according to an embodiment of the present invention is a system that personalizes sports viewing and provides a new entertainment experience. This system allows users to adjust the quality, volume, color, and intensity of the commentator's voice while watching sports. For example, the volume of the commentator can be reduced, and the audio can be set to focus on the atmosphere of the venue. This allows users to enjoy watching sports in an audio environment tailored to their preferences. Next, users can request commentary that focuses on their favorite players. For example, by having the commentary focus on scenes where a specific player is playing, users can enjoy that player's performance more deeply. Furthermore, users can use a communication function with the commentator to converse with the commentator in real time. For example, users can ask questions to the commentator or receive comments from the commentator. This allows users to enjoy two-way communication rather than just passively listening to commentary. In addition, even when watching alone, users can watch while engaging in conversations as if they were watching with multiple people. For example, by watching while conversing with multiple virtual spectators generated by AI, users can have an experience as if they were watching with friends. This system can prevent viewers from dropping out due to commentators, attract new audiences, and increase the overall value of sports. For example, users who find commentators' voices too loud can adjust the volume for a more comfortable viewing experience. Furthermore, providing commentary focused on specific players can attract fans of those players. Additionally, offering a group viewing experience allows users who feel lonely watching alone to enjoy themselves. In this way, the present invention realizes a new entertainment experience for sports viewing by providing a personalized sports viewing experience tailored to user needs. Thus, the sports viewing personalization system can provide a personalized sports viewing experience tailored to user needs.
[0073] The sports viewing personalization system according to this embodiment comprises a reception unit, an adjustment unit, a focus unit, a communication unit, and a provision unit. The reception unit receives user input. User input includes, but is not limited to, voice input, text input, and gesture input. The reception unit analyzes the user's voice input using, for example, speech recognition technology and converts it into text data. The reception unit can also directly receive text input. Furthermore, the reception unit can analyze the user's gesture input using gesture recognition technology and execute a corresponding command. For example, if the user instructs the reception unit by voice to "lower the volume of the commentary," the reception unit analyzes the voice and generates a command to lower the volume of the commentary. The adjustment unit adjusts the quality, quantity, color, and intensity of the commentator's voice based on the information received by the reception unit. The adjustment unit adjusts the tone and volume of the commentator's voice using, for example, voice filtering technology. The adjustment unit can also add echo and reverb to the commentator's voice using effect technology. Furthermore, the adjustment unit can adjust the emotion of the commentator's voice using emotion expression technology. For example, the adjustment unit adjusts the commentator's voice tone to a calmer one, allowing users to relax while watching the game. The focus unit provides commentary focused on specific players based on the information adjusted by the adjustment unit. For example, the focus unit analyzes player performance data and provides commentary centered on scenes where a particular player is playing. The focus unit can also provide commentary focused on the user's favorite players based on user preference information. Furthermore, the focus unit can detect important moments in the match and provide commentary focused on those moments. For example, the focus unit can focus on the moment a particular player scores a goal and provide a detailed explanation of that play. The communication unit communicates with the commentator based on the commentary provided by the focus unit. For example, the communication unit uses voice dialogue technology to enable real-time conversations between the user and the commentator. The communication unit can also use text chat technology to enable text-based communication between the user and the commentator.Furthermore, the communication unit can also use gesture recognition technology to convey the user's gestures to the commentator. For example, the communication unit can have the user ask a question to the commentator and provide the commentator's answer to that question in real time. The delivery unit provides the results of the communication conducted by the communication unit. The delivery unit provides the user with, for example, adjusted audio or generated virtual spectator conversations. The delivery unit can also deliver information to the user's device in real time. Furthermore, the delivery unit can provide customized information tailored to the user's preferences. For example, the delivery unit can provide information about a specific player requested by the user. As a result, the sports viewing personalization system according to the embodiment can adjust the quality, volume, color, and intensity of the commentator's voice based on the user's input, provide commentary focused on a specific player, and communicate with the commentator.
[0074] The reception unit receives user input. User input includes, but is not limited to, voice input, text input, and gesture input. For example, the reception unit analyzes the user's voice input using speech recognition technology and converts it into text data. Specifically, speech recognition technology analyzes the user's speech in real time and converts the speech waveform into a digital signal. Then, it converts the speech data into text data using an acoustic model and a language model. The acoustic model extracts the features of the speech, and the language model selects the most appropriate words and phrases based on the context. This ensures that the user's voice input is accurately converted into text data. The reception unit can also directly accept text input. The user enters text using a keyboard or touchscreen, and that text data is sent to the system. Furthermore, the reception unit can analyze the user's gesture input using gesture recognition technology and execute corresponding commands. Gesture recognition technology detects the user's movements using cameras and sensors and analyzes the movement patterns. For example, if the user waves their hand, that movement is set to correspond to a specific command. This allows the user to operate the system intuitively. For example, if a user gives a voice command saying, "Lower the volume of the commentary," the reception desk analyzes the voice and generates a command to lower the volume. This allows users to easily operate the system and obtain a viewing experience tailored to their preferences.
[0075] The adjustment unit adjusts the quality, volume, tone, and intensity of the commentator's voice based on the information received by the reception unit. For example, the adjustment unit can adjust the tone and volume of the commentator's voice using voice filtering technology. Voice filtering technology improves the quality of the sound by emphasizing or attenuating specific frequency bands. For example, emphasizing low frequency bands can give the commentator's voice more depth. The adjustment unit can also add echo and reverb to the commentator's voice using effects technology. Echo and reverb give the sound a sense of spatial breadth and enhance the feeling of presence. Furthermore, the adjustment unit can adjust the emotion in the commentator's voice using emotion expression technology. Emotion expression technology adds emotion to the commentator's voice by adjusting the pitch, tempo, and volume of the voice. For example, the tone of the commentator's voice can be changed to a calmer tone so that the user can watch the game in a relaxed manner. In this way, the adjustment unit can customize the commentator's voice to suit the user's preferences and provide a more personalized viewing experience.
[0076] The focus unit provides commentary focused on specific players based on information adjusted by the adjustment unit. For example, the focus unit analyzes player performance data and provides commentary centered on the moments when a particular player is playing. Performance data includes player distance covered, number of sprints, and shooting success rate. This data is analyzed in real time to evaluate player performance. For example, it can focus on the moment a particular player scores a goal and provide a detailed explanation of that play. The focus unit can also provide commentary focused on the user's favorite players based on user preference information. User preference information is collected based on past viewing history and user input. This allows users to watch the game focusing on the play of their favorite players. Furthermore, the focus unit can detect important moments in the match and provide commentary focused on those moments. Important moments are detected based on events such as goals, assists, and fouls. This allows users to enjoy important moments without missing the highlights of the match.
[0077] The Communications Department communicates with commentators based on the live commentary provided by the Focus Department. The Communications Department, for example, uses voice dialogue technology to enable real-time conversations between users and commentators. This voice dialogue technology analyzes the user's voice input and generates appropriate responses. For example, if a user asks, "What are this player's past statistics?", the Communications Department analyzes the question, and the commentator provides information about the player's past statistics. The Communications Department can also use text chat technology for text-based communication between users and commentators. This technology analyzes the text entered by the user and generates appropriate responses. Furthermore, the Communications Department can use gesture recognition technology to convey the user's gestures to the commentators. For example, if a user raises their hand, this gesture is interpreted as indicating the intent of a question, and the commentator provides an answer. This allows users to enjoy interactive communication with commentators.
[0078] The service provider provides the results of communication conducted by the communication provider. For example, the service provider provides users with adjusted audio and conversations with generated virtual spectators. Virtual spectators are generated using AI and can enjoy watching the game together with the user. Virtual spectators provide appropriate comments and cheers in response to the user's reactions. This allows users to have an experience as if they were watching with friends, even when watching alone. The service provider can also deliver information to the user's device in real time. For example, it can deliver the latest match information and commentary in real time to smartphones and tablets. Furthermore, the service provider can provide customized information tailored to the user's preferences. For example, it can provide information about a specific player that the user has requested. This allows users to obtain information that matches their interests and enjoy a more fulfilling viewing experience.
[0079] The system includes a virtual spectator unit that generates conversations with virtual spectators. The virtual spectator unit can generate virtual spectators using, for example, AI characters. The virtual spectator unit can also generate the actions and conversations of virtual spectators using simulation models. For example, the virtual spectator unit can generate scenarios where an AI character converses with the user while watching the game. Furthermore, the virtual spectator unit can use simulation models to provide an experience as if multiple virtual spectators were watching the game with the user. For example, the virtual spectator unit can have an AI character generate conversations in real time according to the progress of the match. This allows the user to experience watching the game with multiple people by generating conversations with virtual spectators.
[0080] The adjustment unit can reduce the volume of commentary and set the audio to primarily reflect the atmosphere of the venue. For example, the adjustment unit can reduce the volume of commentary and emphasize crowd cheers and ambient sounds. The adjustment unit can also use audio filtering technology to realistically reproduce the atmosphere of the venue. For example, by emphasizing crowd cheers and reducing the volume of commentary, the adjustment unit can provide users with a sense of presence as if they were actually at the venue. In this way, by reducing the volume of commentary and setting the audio to primarily reflect the atmosphere of the venue, users can enjoy watching sports in an audio environment tailored to their preferences.
[0081] The focus unit can provide commentary centered on moments in which a specific player is playing. For example, the focus unit can analyze player performance data to detect moments in which a particular player is playing. Furthermore, based on user preference information, the focus unit can provide commentary focused on the user's favorite players. For instance, the focus unit can focus on the moment a specific player scores a goal and provide a detailed commentary on that play. This allows users to enjoy that player's performance more deeply by focusing the commentary on moments in which they are playing.
[0082] The communication section allows users to ask questions to commentators and receive comments from them. The communication section enables real-time conversations between users and commentators using, for example, voice dialogue technology. It can also facilitate text-based communication between users and commentators using text chat technology. For instance, the communication section allows users to ask questions to commentators and provides real-time responses from the commentators. This allows users to enjoy two-way communication by asking questions and receiving comments from commentators.
[0083] The service provider can offer users pre-tuned audio and generated virtual spectator conversations. For example, the service provider can deliver pre-tuned audio to the user's device in real time. The service provider can also offer users generated virtual spectator conversations. For example, the service provider can provide information about a specific player requested by the user. By providing users with pre-tuned audio and generated virtual spectator conversations, users can enjoy watching sports tailored to their preferences.
[0084] The reception unit can estimate the user's emotions and adjust the timing of input acceptance based on the estimated emotions. For example, the reception unit can analyze the user's facial expressions using facial recognition technology to estimate emotions. It can also analyze the tone and speed of the user's voice using voice analysis technology to estimate emotions. Furthermore, the reception unit can collect biometric data (heart rate and skin electrical activity) with sensors and estimate emotions using emotion estimation algorithms. For example, if the user is excited, the reception unit can speed up the timing of input acceptance to allow for an immediate response. Conversely, if the user is relaxed, the reception unit can slow down the timing of input acceptance to allow for more time to enter information. Furthermore, if the user is stressed, the reception unit can adjust the timing of input acceptance and provide guidance to reduce stress. In this way, by adjusting the timing of input acceptance based on the user's emotions, input can be accepted at an appropriate time according to the user's emotions. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) and multimodal generation AI.
[0085] The reception desk can analyze a user's past input history and select the optimal input method. For example, the reception desk can prioritize suggesting input methods (voice, text, etc.) that the user has frequently used in the past. It can also predict and suggest input methods to be used during specific time periods based on the user's past input history. Furthermore, the reception desk can provide an auto-completion function to reduce input effort by referencing the user's past input. For example, if the reception desk has frequently used voice input in the past, it will prioritize suggesting voice input. Similarly, if the reception desk has used text input during specific time periods in the past, it can suggest text input during those times. Finally, the reception desk can provide an auto-completion function based on the user's past input to reduce input effort. In this way, by analyzing the user's past input history, the optimal input method can be selected, improving user convenience.
[0086] The reception system can filter input based on the user's current viewing status and areas of interest. For example, if a user is watching a particular sport, the reception system will prioritize receiving only input related to that sport. Similarly, if a user is interested in a particular player, the reception system can prioritize receiving information related to that player. Furthermore, if a user is watching a particular match, the reception system can prioritize receiving only input related to that match. For example, if a user is watching a soccer match, the reception system will prioritize receiving input related to soccer. Similarly, if a user is interested in a particular player, the reception system can prioritize receiving information related to that player. Furthermore, if a user is watching a particular match, the reception system can prioritize receiving input related to that match. This allows the system to prioritize receiving highly relevant information by filtering based on the user's current viewing status and areas of interest.
[0087] The reception unit can estimate the user's emotions and determine the priority of inputs to be received based on the estimated emotions. For example, the reception unit can analyze the user's facial expressions using facial recognition technology to estimate emotions. It can also analyze the tone and speed of the user's voice using voice analysis technology to estimate emotions. Furthermore, the reception unit can collect biometric data (heart rate and skin electrical activity) using sensors and estimate emotions using emotion estimation algorithms. For example, if the user is excited, the reception unit will prioritize receiving important inputs. If the user is relaxed, the reception unit can also prioritize receiving detailed inputs. Furthermore, if the user is stressed, the reception unit can also prioritize receiving simple inputs. In this way, important inputs can be prioritized by determining the priority of inputs based on the user's emotions. Emotion estimation is achieved using emotion estimation functions, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0088] The reception desk can prioritize receiving highly relevant input by considering the user's geographical location. For example, if the user is in a specific region, the reception desk will prioritize receiving information related to that region. Similarly, if the user is in a specific stadium, the reception desk can prioritize receiving information related to that stadium. Furthermore, if the user is in a specific city, the reception desk can prioritize receiving information related to that city. This allows the reception desk to prioritize receiving highly relevant information by considering the user's geographical location.
[0089] The reception desk can analyze the user's social media activity when receiving input and accept relevant input. For example, if the reception desk mentions a specific athlete on social media, it will prioritize receiving information related to that athlete. Similarly, if the reception desk mentions a specific match on social media, it can prioritize receiving information related to that match. Furthermore, if the reception desk mentions a specific sport on social media, it can prioritize receiving information related to that sport. This allows the reception desk to prioritize receiving relevant information by analyzing the user's social media activity.
[0090] The adjustment unit can estimate the user's emotions and adjust the quality, quantity, color, and intensity of the commentary based on the estimated emotions. For example, the adjustment unit can analyze the user's facial expressions using facial recognition technology to estimate emotions. It can also analyze the tone and speed of the user's voice using voice analysis technology to estimate emotions. Furthermore, the adjustment unit can collect biometric data (heart rate and skin electrical activity) using sensors and estimate emotions using emotion estimation algorithms. For example, if the user is excited, the adjustment unit can increase the intensity of the commentary to provide an energetic commentary. If the user is relaxed, the adjustment unit can adjust the quality of the commentary to a calmer tone. Furthermore, if the user is stressed, the adjustment unit can reduce the amount of commentary to provide a simpler commentary. In this way, by adjusting the quality, quantity, color, and intensity of the commentary based on the user's emotions, it is possible to provide commentary that is appropriate to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) and multimodal generation AI.
[0091] The adjustment unit can select the optimal adjustment method by referring to the user's past viewing history during the adjustment process. For example, the adjustment unit can select the optimal adjustment method based on the commentary style the user has preferred in the past. Furthermore, the adjustment unit can prioritize the style of a specific commentator based on the user's past viewing history. In addition, the adjustment unit can analyze the user's past viewing history and select the most appropriate quality and quantity of commentary. This allows the system to select the optimal adjustment method by referring to the user's past viewing history, thereby improving user convenience.
[0092] The adjustment unit can customize the adjustment methods based on the user's current viewing situation. For example, if the user is watching a specific match, the adjustment unit can customize the quality and quantity of commentary to be optimal for that match. It can also customize commentary related to a specific player if the user is focusing on that player. Furthermore, if the user is watching a specific sport, the adjustment unit can customize the quality and quantity of commentary to be optimal for that sport. This allows the system to provide optimal commentary by customizing the adjustment methods based on the user's current viewing situation.
[0093] The adjustment unit can estimate the user's emotions and determine the priority of the commentary to be adjusted based on the estimated emotions. For example, the adjustment unit can analyze the user's facial expressions using facial recognition technology to estimate emotions. It can also analyze the tone and speed of the user's voice using voice analysis technology to estimate emotions. Furthermore, the adjustment unit can collect biometric data (heart rate and skin electrical activity) using sensors and estimate emotions using emotion estimation algorithms. For example, if the user is excited, the adjustment unit can prioritize providing important commentary. It can also prioritize providing detailed commentary if the user is relaxed. Furthermore, it can prioritize providing simple commentary if the user is stressed. In this way, by determining the priority of commentary based on the user's emotions, important commentary can be prioritized. Emotion estimation is achieved using emotion estimation functions, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0094] The adjustment unit can select the optimal adjustment method during adjustment, taking into account the user's geographical location information. For example, if the user is in a specific region, the adjustment unit can prioritize providing explanations related to that region. It can also prioritize providing explanations related to a specific stadium if the user is in that stadium. Furthermore, if the user is in a specific city, the adjustment unit can prioritize providing explanations related to that city. This allows the system to select the optimal adjustment method and improve user convenience by considering the user's geographical location information.
[0095] The adjustment unit can analyze the user's social media activity during the adjustment process and propose adjustment methods. For example, if the user mentions a specific player on social media, the adjustment unit can prioritize providing commentary related to that player. Similarly, if the user mentions a specific match on social media, the adjustment unit can prioritize providing commentary related to that match. Furthermore, if the user mentions a specific sport on social media, the adjustment unit can prioritize providing commentary related to that sport. This allows the system to analyze the user's social media activity, propose the most suitable adjustment method, and improve user convenience.
[0096] The focus unit can estimate the user's emotions and select players or scenes to focus on based on those estimated emotions. For example, the focus unit can analyze the user's facial expressions using facial recognition technology to estimate emotions. It can also analyze the tone and speed of the user's voice using voice analysis technology to estimate emotions. Furthermore, the focus unit can collect biometric data (heart rate and skin electrical activity) using sensors and estimate emotions using emotion estimation algorithms. For example, if the user is excited, the focus unit will prioritize focusing on important players or scenes. If the user is relaxed, the focus unit can prioritize focusing on detailed players or scenes. Furthermore, if the user is stressed, the focus unit can prioritize focusing on simple players or scenes. This allows for optimal focus tailored to the user's emotions by selecting players or scenes based on their feelings. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0097] The focus unit can select the optimal focus method by referring to the user's past viewing history when focusing. For example, the focus unit can select the optimal focus method based on the player or scene the user has liked in the past. The focus unit can also prioritize focusing on specific players or scenes based on the user's past viewing history. Furthermore, the focus unit can analyze the user's past viewing history and focus on the most suitable player or scene. For example, the focus unit can select the optimal focus method based on the player or scene the user has liked in the past. Furthermore, the focus unit can prioritize focusing on specific players or scenes based on the user's past viewing history. Furthermore, the focus unit can analyze the user's past viewing history and focus on the most suitable player or scene. By referring to the user's past viewing history, the optimal focus method can be selected, improving user convenience.
[0098] The focus unit can customize the means of focusing based on the user's current viewing situation. For example, if the user is watching a specific match, the focus unit will focus on the most suitable players or scenes for that match. It can also focus on scenes related to a specific player if the user is focusing on that player. Furthermore, if the user is watching a specific sport, the focus unit can focus on the most suitable players or scenes for that sport. This allows the system to provide optimal focus by customizing the means of focusing based on the user's current viewing situation.
[0099] The focus unit can estimate the user's emotions and determine the priority of players or scenes to focus on based on the estimated emotions. For example, the focus unit can analyze the user's facial expressions using facial recognition technology to estimate emotions. It can also analyze the tone and speed of the user's voice using voice analysis technology to estimate emotions. Furthermore, the focus unit can collect biometric data (heart rate and skin electrical activity) using sensors and estimate emotions using emotion estimation algorithms. For example, if the user is excited, the focus unit will prioritize focusing on important players or scenes. If the user is relaxed, the focus unit can also prioritize focusing on detailed players or scenes. Furthermore, if the user is stressed, the focus unit can also prioritize focusing on simple players or scenes. In this way, by determining the priority of focus based on the user's emotions, it is possible to prioritize focusing on important players and scenes. Emotion estimation is achieved using emotion estimation functions, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0100] The focus unit can select the optimal focus method when focusing, taking into account the user's geographical location. For example, if the user is in a specific region, the focus unit will prioritize focusing on players or scenes related to that region. Similarly, if the user is in a specific stadium, the focus unit can prioritize focusing on players or scenes related to that stadium. Furthermore, if the user is in a specific city, the focus unit can prioritize focusing on players or scenes related to that city. This allows the system to select the optimal focus method by considering the user's geographical location, thereby improving user convenience.
[0101] The focus unit can analyze the user's social media activity and suggest a method of focus when focusing. For example, if the user mentions a specific athlete on social media, the focus unit will prioritize focusing on scenes related to that athlete. Similarly, if the user mentions a specific match on social media, the focus unit can prioritize focusing on scenes related to that match. Furthermore, if the user mentions a specific sport on social media, the focus unit can prioritize focusing on scenes related to that sport. This allows the system to analyze the user's social media activity, suggest the optimal method of focus, and improve user convenience.
[0102] The communication unit can estimate the user's emotions and adjust its communication method based on those emotions. For example, the communication unit can analyze the user's facial expressions using facial recognition technology to estimate emotions. It can also analyze the tone and speed of the user's voice using voice analysis technology to estimate emotions. Furthermore, the communication unit can collect biometric data (heart rate and skin electrical activity) using sensors and estimate emotions using emotion estimation algorithms. For example, if the user is excited, the communication unit will communicate in an energetic tone. If the user is relaxed, the communication unit can communicate in a calm tone. Furthermore, if the user is stressed, the communication unit can communicate in a gentle tone. In this way, by adjusting the communication method based on the user's emotions, it is possible to provide optimal communication that is appropriate to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0103] The communication department can select the optimal method of communication by referring to the user's past communication history. For example, the communication department can select the optimal method based on the communication style the user has preferred in the past. Furthermore, the communication department can prioritize specific tones and styles based on the user's past communication history. In addition, the communication department can analyze the user's past communication history and select the most appropriate communication method. This allows the system to select the optimal communication method by referring to the user's past communication history, thereby improving user convenience.
[0104] The communications department can customize the means of communication based on the user's current viewing situation. For example, if the user is watching a particular match, the communications department can customize the communication method best suited to that match. Furthermore, if the user is focusing on a particular player, the communications department can customize communication related to that player. Additionally, if the user is watching a particular sport, the communications department can customize communication methods best suited to that sport. This allows for optimal communication by customizing the means of communication based on the user's current viewing situation.
[0105] The communication unit can estimate the user's emotions and determine communication priorities based on those estimated emotions. For example, the communication unit can analyze the user's facial expressions using facial recognition technology to estimate emotions. It can also analyze the tone and speed of the user's voice using voice analysis technology to estimate emotions. Furthermore, the communication unit can collect biometric data (heart rate and skin electrical activity) using sensors and estimate emotions using emotion estimation algorithms. For example, if the user is excited, the communication unit will prioritize important communication. If the user is relaxed, it will prioritize detailed communication. Furthermore, if the user is stressed, it will prioritize simple communication. This allows for prioritizing important communication based on the user's emotions. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0106] The communications department can select the optimal method of communication by considering the user's geographical location. For example, if the user is in a specific region, the communications department can prioritize communications related to that region. Similarly, if the user is in a specific stadium, the communications department can prioritize communications related to that stadium. Furthermore, if the user is in a specific city, the communications department can prioritize communications related to that city. This allows the communications department to select the optimal communication method by considering the user's geographical location, thereby improving user convenience.
[0107] The communications department can analyze users' social media activity and suggest appropriate communication methods during communication. For example, if a user mentions a specific athlete on social media, the communications department will prioritize communications related to that athlete. Similarly, if a user mentions a specific match on social media, the communications department can prioritize communications related to that match. Furthermore, if a user mentions a specific sport on social media, the communications department can prioritize communications related to that sport. By analyzing users' social media activity, the communications department can suggest optimal communication methods and improve user convenience.
[0108] The information provider can estimate the user's emotions and adjust the format of the information provided based on the estimated emotions. For example, the provider can analyze the user's facial expressions using facial recognition technology to estimate emotions. It can also analyze the tone and speed of the user's voice using voice analysis technology to estimate emotions. Furthermore, the provider can collect biometric data (heart rate and skin electrical activity) using sensors and estimate emotions using emotion estimation algorithms. For example, if the user is excited, the provider can provide information in a visually stimulating format. If the user is relaxed, the provider can provide information in a calm format. Furthermore, if the user is stressed, the provider can provide information in a simple and highly visible format. This allows the provider to provide optimal information tailored to the user's emotions by adjusting the format of the information based on those emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0109] The service provider can select the optimal delivery method by referring to the user's past viewing history at the time of delivery. For example, the service provider can select the optimal delivery method based on the information format the user has preferred in the past. Furthermore, the service provider can prioritize the selection of a specific information format based on the user's past viewing history. In addition, the service provider can analyze the user's past viewing history and select the most suitable information format. This allows the service provider to select the optimal delivery method by referring to the user's past viewing history, thereby improving user convenience.
[0110] The service provider can customize the means of delivery based on the user's current viewing situation. For example, if the user is watching a particular match, the service provider can customize the information format to be best suited to that match. Furthermore, if the user is focusing on a particular player, the service provider can customize the information format to be best suited to that player. In addition, if the user is watching a particular sport, the service provider can customize the information format to be best suited to that sport. This allows the service provider to provide optimal information by customizing the means of delivery based on the user's current viewing situation.
[0111] The information provider can estimate the user's emotions and prioritize the information to be provided based on those emotions. For example, the provider can analyze the user's facial expressions using facial recognition technology to estimate emotions. It can also analyze the tone and speed of the user's voice using voice analysis technology to estimate emotions. Furthermore, the provider can collect biometric data (heart rate and skin electrical activity) using sensors and estimate emotions using emotion estimation algorithms. For example, if the user is excited, the provider will prioritize providing important information. If the user is relaxed, the provider will also prioritize providing detailed information. Furthermore, if the user is stressed, the provider will prioritize providing simple information. This allows for the prioritization of important information based on the user's emotions. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0112] The service provider can select the optimal delivery method by considering the user's geographical location at the time of delivery. For example, if the user is in a specific region, the service provider can prioritize providing information related to that region. Similarly, if the user is in a specific stadium, the service provider can prioritize providing information related to that stadium. Furthermore, if the user is in a specific city, the service provider can prioritize providing information related to that city. This allows the service provider to select the optimal delivery method by considering the user's geographical location, thereby improving user convenience.
[0113] The service provider can analyze the user's social media activity at the time of delivery and propose a delivery method. For example, if the service provider mentions a specific athlete on social media, it can prioritize providing information related to that athlete. Similarly, if the service provider mentions a specific match on social media, it can prioritize providing information related to that match. Furthermore, if the service provider mentions a specific sport on social media, it can prioritize providing information related to that sport. By analyzing the user's social media activity, the service provider can propose the optimal delivery method and improve user convenience.
[0114] The virtual spectator unit can estimate the user's emotions and adjust the content of the virtual spectator's conversation based on the estimated emotions. For example, the virtual spectator unit can analyze the user's facial expressions using facial recognition technology to estimate emotions. It can also analyze the tone and speed of the user's voice using voice analysis technology to estimate emotions. Furthermore, the virtual spectator unit can collect biometric data (heart rate and skin electrical activity) using sensors and estimate emotions using emotion estimation algorithms. For example, if the virtual spectator unit is excited, it will provide energetic conversation content. If the user is relaxed, it will provide calm conversation content. Furthermore, if the user is stressed, it will provide gentle conversation content. In this way, by adjusting the content of the virtual spectator's conversation based on the user's emotions, it is possible to provide optimal conversation that matches the user's emotions. Emotion estimation is achieved using emotion estimation functions, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.
[0115] The virtual spectator unit can select the most suitable conversation content by referring to the user's past viewing history when generating conversations for virtual spectators. For example, the virtual spectator unit can select the most suitable conversation content based on the conversation content the user has enjoyed in the past. Furthermore, the virtual spectator unit can prioritize the selection of specific conversation content from the user's past viewing history. In addition, the virtual spectator unit can analyze the user's past viewing history and select the most suitable conversation content. This allows for the selection of the most suitable conversation content by referring to the user's past viewing history, thereby improving user convenience.
[0116] The virtual spectator unit can customize conversation content based on the user's current viewing situation when generating conversations with virtual spectators. For example, if the user is watching a specific match, the virtual spectator unit will customize conversation content related to that match. It can also customize conversation content related to a specific player if the user is focusing on that player. Furthermore, if the virtual spectator unit is watching a specific sport, it can customize conversation content related to that sport. This allows the system to provide optimal conversations by customizing conversation content based on the user's current viewing situation.
[0117] The virtual spectator unit can estimate the user's emotions and prioritize conversations based on those emotions. For example, the virtual spectator unit can analyze the user's facial expressions using facial recognition technology to estimate emotions. It can also analyze the tone and speed of the user's voice using voice analysis technology to estimate emotions. Furthermore, the virtual spectator unit can collect biometric data (heart rate and skin electrical activity) using sensors and estimate emotions using emotion estimation algorithms. For example, if the virtual spectator unit is excited, it will prioritize important conversations. If the user is relaxed, it will prioritize detailed conversations. Furthermore, if the user is stressed, it will prioritize simple conversations. This allows for prioritizing important conversations based on the user's emotions. Emotion estimation is achieved using emotion estimation functions, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) and multimodal generation AI.
[0118] The virtual spectator system can select the most relevant conversation content when generating conversations for virtual spectators, taking into account the user's geographical location. For example, if the user is in a specific region, the system will prioritize providing conversation content related to that region. It can also prioritize providing conversation content related to a specific stadium if the user is in that stadium. Furthermore, if the user is in a specific city, the system can prioritize providing conversation content related to that city. This allows the system to select the most relevant conversation content by considering the user's geographical location, thereby improving user convenience.
[0119] The virtual spectator system can analyze a user's social media activity to select the most relevant conversation content when generating conversations for virtual spectators. For example, if a user mentions a specific player on social media, the system will prioritize providing conversation content related to that player. Similarly, if a user mentions a specific match on social media, the system can prioritize providing conversation content related to that match. Furthermore, if a user mentions a specific sport on social media, the system can prioritize providing conversation content related to that sport. This allows the system to analyze a user's social media activity, select the most relevant conversation content, and improve user experience.
[0120] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0121] A personalized sports viewing system can monitor a user's health and adjust the viewing experience accordingly. For example, it can measure heart rate and blood pressure with sensors and, if the user is overly excited, lower the volume of the commentary or play relaxing music. If the user is tired, it can change the tone of the commentary to a calmer one and reduce visual stimulation. Furthermore, if the user has health problems, it can display appropriate alerts and encourage them to take a break. This allows for the provision of an optimal viewing experience tailored to the user's health condition.
[0122] A personalized sports viewing system can provide customized content based on the user's hobbies and interests. For example, if a user is interested in movies or music in addition to a particular sport, the system can provide information and content related to those interests while they are watching a sport. Similarly, if a user is interested in a specific region or culture, the system can provide information related to that region or culture. Furthermore, if a user is interested in a specific event or festival, the system can provide information related to that event. This allows for a customized viewing experience based on the user's hobbies and interests.
[0123] The personalized sports viewing system can estimate a user's emotions and adjust the advertisements displayed during the event based on those emotions. For example, if a user is excited, it can display energetic advertisements to further enhance their excitement. If a user is relaxed, it can display advertisements with a calm tone to maintain a relaxed atmosphere. Furthermore, if a user is stressed, it can display relaxation advertisements to alleviate stress. This allows the system to provide optimal advertisements tailored to the user's emotions.
[0124] A personalized sports viewing system can analyze a user's viewing history and provide information related to their favorite matches and players in the past. For example, it can provide highlights of matches the user has watched in the past or compilations of specific players' plays. It can also suggest the most suitable commentator based on the user's preferred commentary style and announcers. Furthermore, it can provide information related to events and festivals the user has attended in the past. This allows for a customized viewing experience based on the user's viewing history.
[0125] A personalized sports viewing system can estimate a user's emotions and adjust interactive elements during viewing based on those emotions. For example, if a user is excited, it can offer interactive games or quizzes to further enhance their excitement. If a user is relaxed, it can offer relaxing interactive elements to maintain a relaxed atmosphere. Furthermore, if a user is stressed, it can offer relaxation games or quizzes to alleviate stress. This allows the system to provide optimal interactive elements tailored to the user's emotions.
[0126] A personalized sports viewing system can customize the information provided during a game by taking into account the user's geographical location. For example, if the user is in a specific region, it can provide sports news and event information related to that region. If the user is at a specific stadium, it can provide information and history related to that stadium. Furthermore, if the user is in a specific city, it can provide tourist and cultural information related to that city. This allows for a customized viewing experience based on the user's geographical location.
[0127] The personalized sports viewing system can estimate the user's emotions and adjust the music played during the event based on those emotions. For example, if the user is excited, it can play energetic music to further enhance their excitement. If the user is relaxed, it can play calming music to maintain a relaxed atmosphere. Furthermore, if the user is stressed, it can play relaxation music to reduce stress. This allows the system to provide optimal music tailored to the user's emotions.
[0128] A personalized sports viewing system can analyze a user's social media activity and customize the information provided during viewing. For example, if a user mentions a specific player on social media, it can prioritize providing information related to that player. Similarly, if a user mentions a specific match on social media, it can prioritize providing information related to that match. Furthermore, if a user mentions a specific sport on social media, it can prioritize providing information related to that sport. This allows for a customized viewing experience based on the user's social media activity.
[0129] The personalized sports viewing system can estimate the user's emotions and adjust the visual effects during viewing based on those emotions. For example, if the user is excited, the system can emphasize the visual effects to further enhance their excitement. If the user is relaxed, the system can change the visual effects to calmer ones to maintain a relaxed atmosphere. Furthermore, if the user is stressed, the system can simplify the visual effects and reduce visual stimulation. This allows the system to provide optimal visual effects tailored to the user's emotions.
[0130] A personalized sports viewing system can monitor a user's behavior while watching a game and adjust the viewing experience accordingly. For example, if a user frequently gets up from their seat, it can display alerts to ensure they don't miss important moments. It can also display alerts encouraging users to take breaks if they are sitting in the same position for extended periods. Furthermore, if a user repeatedly engages in certain behaviors while watching, the system can provide information and content related to those behaviors. This allows for a customized viewing experience based on the user's actions during the game.
[0131] The following briefly describes the processing flow for example form 2.
[0132] Step 1: The reception unit receives user input. User input includes voice input, text input, and gesture input. For example, voice recognition technology can be used to analyze the user's voice input and convert it into text data. It can also directly accept text input, and gesture recognition technology can be used to analyze the user's gesture input and execute the corresponding command. Step 2: The adjustment unit adjusts the quality, volume, color, and intensity of the commentator's voice based on the information received by the reception unit. For example, it can use voice filtering technology to adjust the tone and volume of the commentator's voice, and effect technology to add echo and reverb to the commentator's voice. Furthermore, it can also adjust the emotion of the commentator's voice using emotion expression technology. Step 3: The focus unit provides commentary focused on a specific player based on the information adjusted by the adjustment unit. For example, it can analyze player performance data and provide commentary focusing on scenes where a particular player is playing. It can also provide commentary focused on the user's favorite player based on user preference information. Furthermore, it can detect important moments in the match and provide commentary focused on those moments. Step 4: The communication unit communicates with the commentator based on the live commentary provided by the focus unit. For example, voice dialogue technology can be used to enable real-time conversations between the user and the commentator, and text chat technology can be used for text-based communication between the user and the commentator. Furthermore, gesture recognition technology can be used to convey the user's gestures to the commentator. Step 5: The delivery unit provides the results of the communication conducted by the communication unit. For example, it can provide users with coordinated audio or generated virtual spectator conversations, and can also deliver information to the user's device in real time. Furthermore, it can provide customized information tailored to the user's preferences.
[0133] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0134] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0135] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0136] Each of the multiple elements described above, including the reception unit, adjustment unit, focus unit, communication unit, provision unit, and virtual viewing unit, is implemented, for example, by at least one of the smart device 14 and the data processing unit 12. For example, the reception unit receives user input using the microphone 38B and touch panel 38A of the smart device 14. The adjustment unit adjusts the quality and volume of the commentator's voice using the specific processing unit 290 of the data processing unit 12. The focus unit provides commentary focused on a specific player using the specific processing unit 290 of the data processing unit 12. The communication unit enables real-time conversation with the commentator using the control unit 46A of the smart device 14. The provision unit provides the user with adjusted audio and conversations with virtual viewers using the output device 40 of the smart device 14. The virtual viewing unit generates an AI character using the specific processing unit 290 of the data processing unit 12 and provides a scenario in which the user can watch the game while conversing with the AI character. The correspondence between each unit and the devices and control units is not limited to the example described above and can be modified in various ways.
[0137] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0138] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0139] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0140] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0141] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0142] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0143] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0144] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0145] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0146] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0147] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0148] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0149] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0150] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0151] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0152] Each of the multiple elements described above, including the reception unit, adjustment unit, focus unit, communication unit, provision unit, and virtual viewing unit, is implemented by, for example, at least one of the smart glasses 214 and the data processing unit 12. For example, the reception unit receives the user's voice input using the microphone 238 of the smart glasses 214. The adjustment unit adjusts the quality and volume of the commentator's voice using the specific processing unit 290 of the data processing unit 12. The focus unit provides commentary focused on a specific player using the specific processing unit 290 of the data processing unit 12. The communication unit enables real-time conversation with the commentator using the control unit 46A of the smart glasses 214. The provision unit provides the user with the adjusted voice and conversations with the virtual viewing unit using the speaker 240 of the smart glasses 214. The virtual viewing unit generates an AI character using the specific processing unit 290 of the data processing unit 12 and provides a scenario in which the user can watch the game while conversing with the AI character. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.
[0153] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0154] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0155] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0156] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0157] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0158] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0159] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0160] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0161] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0162] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0163] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0164] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0165] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0166] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0167] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0168] Each of the multiple elements described above, including the reception unit, adjustment unit, focus unit, communication unit, provision unit, and virtual viewing unit, is implemented by, for example, at least one of the headset terminal 314 and the data processing unit 12. For example, the reception unit receives the user's voice input using the microphone 238 of the headset terminal 314. The adjustment unit adjusts the quality and volume of the commentator's voice using the specific processing unit 290 of the data processing unit 12. The focus unit provides commentary focused on a specific player using the specific processing unit 290 of the data processing unit 12. The communication unit enables real-time conversation with the commentator using the control unit 46A of the headset terminal 314. The provision unit provides the user with the adjusted voice and conversations of the virtual viewing unit using the speaker 240 of the headset terminal 314. The virtual viewing unit generates an AI character using the specific processing unit 290 of the data processing unit 12 and provides a scenario in which the user can watch the game while conversing with the AI character. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.
[0169] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0170] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0171] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0172] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0173] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0174] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0175] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0176] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0177] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0178] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0179] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0180] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0181] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0182] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0183] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0184] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0185] Each of the multiple elements described above, including the reception unit, adjustment unit, focus unit, communication unit, provision unit, and virtual viewing unit, is implemented by, for example, at least one of the robot 414 and the data processing unit 12. For example, the reception unit receives voice input from the user using the microphone 238 of the robot 414. The adjustment unit adjusts the quality and volume of the commentator's voice using the specific processing unit 290 of the data processing unit 12. The focus unit provides commentary focused on a specific player using the specific processing unit 290 of the data processing unit 12. The communication unit enables real-time conversation with the commentator using the control unit 46A of the robot 414. The provision unit provides the user with the adjusted voice and conversations of the virtual viewing unit using the speaker 240 of the robot 414. The virtual viewing unit generates an AI character using the specific processing unit 290 of the data processing unit 12 and provides a scenario in which the user can watch the game while conversing with the AI character. The correspondence between each unit and the devices and control units is not limited to the example described above and can be changed in various ways.
[0186] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0187] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0188] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0189] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0190] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0191] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0192] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0193] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0194] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0195] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0196] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0197] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0198] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0199] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0200] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0201] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0202] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0203] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0204] (Note 1) A reception area that receives user input, Based on the information received by the reception unit, an adjustment unit adjusts the quality, volume, color, and intensity of the commentator's voice. A focus unit provides commentary that focuses on a specific player based on the information adjusted by the adjustment unit, A communication unit that communicates with commentators based on the live commentary performed by the aforementioned focus unit, A providing unit that provides the results of communication conducted by the aforementioned communication unit. A system characterized by the following features. (Note 2) It includes a virtual spectator unit that generates conversations with virtual spectators. The system described in Appendix 1, characterized by the features described herein. (Note 3) The adjustment unit is, Reduce the volume of commentary and focus the audio on conveying the atmosphere of the venue. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned focusing unit is The commentary will focus on scenes where specific players are playing. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned communications department, Users can ask questions to commentators and receive comments from them. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned supply unit is, Provides users with adjusted audio and generated virtual spectator conversations. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned reception unit is The system estimates the user's emotions and adjusts the timing of input acceptance based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned reception unit is Analyze the user's past input history and select the optimal input method. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned reception unit is When receiving input, the system filters it based on the user's current viewing status and areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned reception unit is It estimates the user's emotions and determines the priority of input to accept based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned reception unit is When receiving input, the system prioritizes accepting inputs that are highly relevant, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned reception unit is When receiving input, the system analyzes the user's social media activity and accepts relevant input. The system described in Appendix 1, characterized by the features described herein. (Note 13) The adjustment unit is, It estimates the user's emotions and adjusts the quality, quantity, color, and intensity of the explanation based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The adjustment unit is, During adjustments, the system will refer to the user's past viewing history to select the optimal adjustment method. The system described in Appendix 1, characterized by the features described herein. (Note 15) The adjustment unit is, During adjustments, the adjustment method is customized based on the user's current viewing status. The system described in Appendix 1, characterized by the features described herein. (Note 16) The adjustment unit is, It estimates the user's emotions and determines the priority of explanations to adjust based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The adjustment unit is, During the adjustment process, the optimal adjustment method is selected by considering the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 18) The adjustment unit is, During the adjustment process, we analyze users' social media activity and propose adjustment methods. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned focusing unit is The system estimates the user's emotions and selects players or scenes to focus on based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned focusing unit is When focusing, the system selects the optimal focusing method by referring to the user's past viewing history. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned focusing unit is When focusing, the method of focusing is customized based on the user's current viewing status. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned focusing unit is It estimates the user's emotions and, based on those emotions, determines the priority of players and scenes to focus on. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned focusing unit is When focusing, the system selects the optimal focusing method considering the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned focusing unit is During the focus phase, we analyze users' social media activity and suggest methods for focusing. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned communications department, It estimates the user's emotions and adjusts the communication method based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned communications department, During communication, the system selects the optimal method by referring to the user's past communication history. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned communications department, During communication, the means of communication are customized based on the user's current viewing status. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned communications department, It estimates the user's emotions and determines communication priorities based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned communications department, When communicating, the optimal method is selected considering the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned communications department, During communication, we analyze the user's social media activity and suggest communication methods. The system described in Appendix 1, characterized by the features described herein. (Note 31) The aforementioned supply unit is, It estimates the user's emotions and adjusts the format of the information provided based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 32) The aforementioned supply unit is, When providing the service, the system will refer to the user's past viewing history to select the most suitable delivery method. The system described in Appendix 1, characterized by the features described herein. (Note 33) The aforementioned supply unit is, When providing the service, the method of delivery will be customized based on the user's current viewing status. The system described in Appendix 1, characterized by the features described herein. (Note 34) The aforementioned supply unit is, It estimates the user's emotions and prioritizes the information provided based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 35) The aforementioned supply unit is, When providing the service, the optimal delivery method will be selected, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 36) The aforementioned supply unit is, When providing the service, we analyze the user's social media activity and propose a delivery method. The system described in Appendix 1, characterized by the features described herein. (Note 37) The aforementioned virtual spectator section, It estimates the user's emotions and adjusts the conversation content of the virtual spectators based on the estimated user emotions. The system described in Appendix 2, characterized by the features described herein. (Note 38) The aforementioned virtual spectator section, When generating conversations for virtual spectators, the system selects the most appropriate conversation content by referring to the user's past viewing history. The system described in Appendix 2, characterized by the features described herein. (Note 39) The aforementioned virtual spectator section, When generating conversations for virtual spectators, the conversation content is customized based on the user's current viewing status. The system described in Appendix 2, characterized by the features described herein. (Note 40) The aforementioned virtual spectator section, It estimates the user's emotions and determines the priority of virtual spectator conversations based on the estimated user emotions. The system described in Appendix 2, characterized by the features described herein. (Note 41) The aforementioned virtual spectator section, When generating conversations for virtual spectators, the system selects the most appropriate conversation content by considering the user's geographical location. The system described in Appendix 2, characterized by the features described herein. (Note 42) The aforementioned virtual spectator section, When generating conversations for virtual spectators, the system analyzes the user's social media activity to select the most appropriate conversation content. The system described in Appendix 2, characterized by the features described herein. [Explanation of Symbols]
[0205] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. A reception area that receives user input, Based on the information received by the reception unit, an adjustment unit adjusts the quality, volume, color, and intensity of the commentator's voice. A focus unit provides commentary that focuses on a specific player based on the information adjusted by the adjustment unit, A communication unit that communicates with commentators based on the live commentary performed by the aforementioned focus unit, A providing unit that provides the results of communication conducted by the aforementioned communication unit. A system characterized by the following features.
2. It includes a virtual spectator unit that generates conversations with virtual spectators. The system according to feature 1.
3. The adjustment unit is, Reduce the volume of commentary and focus the audio on conveying the atmosphere of the venue. The system according to feature 1.
4. The aforementioned focusing unit is The commentary will focus on scenes where specific players are playing. The system according to feature 1.
5. The aforementioned communications department, Users can ask questions to commentators and receive comments from them. The system according to feature 1.
6. The aforementioned supply unit is, Provides users with adjusted audio and generated virtual spectator conversations. The system according to feature 1.
7. The aforementioned reception unit is The system estimates the user's emotions and adjusts the timing of input acceptance based on the estimated emotions. The system according to feature 1.
8. The aforementioned reception unit is Analyze the user's past input history and select the optimal input method. The system according to feature 1.
9. The aforementioned reception unit is When receiving input, the system filters it based on the user's current viewing status and areas of interest. The system according to feature 1.
10. The aforementioned reception unit is It estimates the user's emotions and determines the priority of input to accept based on the estimated user emotions. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A