system
The system addresses the challenge of dynamic video editing by automating scene extraction and collaborative editing, ensuring viewer preferences are met, thereby improving satisfaction and generating revenue.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Editing of live events or stream videos requires labor and time, and it is difficult to perform dynamic editing according to the preferences of viewers.
A system comprising an extraction unit, an editing unit, a proposal unit, and a distribution unit that automatically extracts scenes from live events or streamed video, allows viewers to collaboratively edit them, and dynamically edits the video to suit viewer preferences before distribution.
Efficiently edits live events and streamed video to meet viewer preferences, enhancing viewer satisfaction and enabling revenue generation through subscriptions and advertising integration.
Smart Images

Figure 2026073113000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there is a problem that editing of live events or stream videos requires labor and time, and it is difficult to perform dynamic editing according to the preferences of viewers.
[0005] The system according to the embodiment aims to efficiently edit live events or stream videos according to the preferences of viewers.
Means for Solving the Problems
[0006] The system according to this embodiment comprises an extraction unit, an editing unit, a proposal unit, and a distribution unit. The extraction unit automatically extracts multiple scenes from a live event or streamed video. The editing unit allows viewers to collaboratively edit the scenes extracted by the extraction unit. The proposal unit dynamically edits the video edited by the editing unit to suit the viewer's preferences. The distribution unit approves the editing results proposed by the proposal unit and distributes them to a streaming platform. [Effects of the Invention]
[0007] The system according to this embodiment can efficiently edit live events and streamed video to suit the viewer's preferences. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10]This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The automated video editing platform according to an embodiment of the present invention is a system that utilizes generative AI and advanced video analysis to automatically extract multiple scenes from live events and streamed video, and allows viewers to collaboratively edit them. This system uses AI to analyze trends and sentiment, and proposes dynamic editing tailored to each viewer's preferences. The collaborative results are delivered to the streaming platform after approval. The main revenue sources are subscriptions to the Pro plan and revenue sharing through advertising partnerships. For example, the AI automatically extracts multiple scenes from live events and streamed video. For instance, it can extract important scenes from live footage of a sports event. Next, viewers can collaboratively edit the video. In this process, viewers can select and edit scenes according to their preferences. Furthermore, the AI analyzes trends and sentiment, and proposes dynamic editing tailored to each viewer's preferences. For example, if a viewer prefers emotionally moving scenes, the AI can prioritize suggesting such scenes. In this way, viewers can edit according to their preferences. The collaborative results are delivered to the streaming platform after viewer approval. This allows viewers to share the videos they participated in editing with other viewers. Furthermore, the main sources of revenue are subscriptions to the Pro plan and revenue sharing through advertising integration. The Pro plan subscription allows viewers to use more advanced editing features. Additionally, advertising integration allows revenue to be generated by inserting advertisements into videos edited by viewers. In this way, the automated video editing platform, utilizing generative AI and advanced video analysis, enables collaborative editing among viewers and increases viewer satisfaction by suggesting dynamic edits tailored to each viewer's preferences. Moreover, the revenue sharing through Pro plan subscriptions and advertising integration enables a sustainable business model. Thus, the automated video editing platform enables collaborative editing among viewers and suggests dynamic edits tailored to each viewer's preferences.
[0029] The automated video editing platform according to this embodiment comprises an extraction unit, an editing unit, a suggestion unit, and a distribution unit. The extraction unit automatically extracts multiple scenes from live events or streamed video. For example, the extraction unit can extract important scenes from live video of a sports event. The extraction unit uses image recognition technology and audio analysis technology to detect specific events or actions in the video and extracts scenes based on them. For example, the extraction unit can automatically detect and extract scoring scenes or highlight scenes. The editing unit allows viewers to collaboratively edit the scenes extracted by the extraction unit. For example, the editing unit allows viewers to select and edit scenes according to their preferences. The editing unit provides online collaboration tools and real-time editing functions to enable viewers to work together on editing. For example, the editing unit allows viewers to change the order of scenes or add effects. The suggestion unit dynamically edits the video edited by the editing unit according to the viewer's preferences. For example, the suggestion unit identifies the viewer's preferences based on the viewer's past viewing history or survey results and makes editing suggestions based on them. The suggestion unit can prioritize suggesting emotionally impactful scenes if viewers prefer them. The suggestion unit uses AI to analyze viewers' preferences and proposes optimal editing. The distribution unit approves the editing results proposed by the suggestion unit and distributes them to the streaming platform. By distributing the editing results after viewer approval, the distribution unit enables viewers to share the video they participated in editing with other viewers. The distribution unit approves and distributes the editing results based on viewer votes and feedback. As a result, the automated video editing platform according to this embodiment can automatically extract multiple scenes from live events and streamed video, allow viewers to collaboratively edit them, propose dynamic editing tailored to viewers' preferences, and distribute it to the streaming platform.
[0030] The extraction unit automatically extracts multiple scenes from live events and streamed video. Specifically, it can extract important scenes from live video of sports events. The extraction unit uses image recognition and audio analysis technologies to detect specific events and actions within the video and extracts scenes based on that. For example, the extraction unit can automatically detect and extract scoring scenes and highlight scenes. Image recognition technology utilizes object detection algorithms based on deep learning to identify specific movements and objects within the video. For example, in sports events, it can track the movements of players and the position of the ball to identify scoring scenes and important plays. Audio analysis technology analyzes audio signals and detects changes in the tone of cheers and commentary to identify important scenes. For example, it can detect moments when the crowd's cheers get louder or when the commentator is excited and extract those scenes. Furthermore, the extraction unit can integrate video from multiple camera angles and select the optimal scene. This allows for the provision of the most engaging video for viewers. Because the extraction unit analyzes video in real time and instantly extracts important scenes, it can respond quickly even during live events. This allows viewers to enjoy the show in real time without missing any important scenes.
[0031] The editorial team allows viewers to collaboratively edit scenes extracted by the extraction team. Specifically, viewers can select and edit scenes to their liking. The editorial team provides online collaboration tools and real-time editing functions, enabling viewers to work together on the editing process. For example, the editorial team allows viewers to change the order of scenes or add effects. Online collaboration tools enable multiple viewers to work on editing simultaneously and share changes in real time. Viewers can use chat and comment functions to exchange opinions with other viewers as they work on the editing process. The real-time editing function allows viewers to instantly check their edits and make corrections as needed. For example, if a viewer changes the order of scenes, the change is immediately reflected, and the viewer can check the editing results in real time. Furthermore, the editorial team provides a variety of effects and transitions for viewers to use, allowing them to customize the video to their liking. This enables viewers to create their own original videos and share them with other viewers. The editorial team can draw out the creativity of viewers and provide them with the enjoyment of video editing.
[0032] The suggestion department dynamically edits the videos edited by the editorial department to match the viewer's preferences. Specifically, it identifies viewer preferences based on their past viewing history and survey results, and makes editing suggestions accordingly. If a viewer likes emotionally moving scenes, the suggestion department can prioritize suggesting such scenes. The suggestion department uses AI to analyze viewer preferences and propose the optimal editing. The AI analyzes the viewer's viewing history and survey results to identify their preferences and interests. For example, if a viewer has watched many emotionally moving scenes in the past, the AI will determine that the viewer likes emotionally moving scenes and will prioritize suggesting such scenes. The AI can also analyze the viewer's real-time reactions to understand their interests and concerns. For example, if a viewer watches a particular scene for a long time, the AI will determine that the scene is interesting to the viewer and can suggest similar scenes. Furthermore, the suggestion department can collect viewer feedback and continuously improve its suggestions. By providing evaluations and comments on the suggested editing, the suggestion department can more accurately understand viewer preferences and reflect them in future suggestions. This allows the proposal department to provide optimal editing tailored to the viewer's preferences, thereby improving the viewing experience.
[0033] The distribution department approves the editing results proposed by the proposal department and distributes them to streaming platforms. Specifically, by distributing the edited results after viewer approval, viewers can share videos they have edited with other viewers. The distribution department approves and distributes the edited results based on viewer votes and feedback. For example, viewers can vote on the edited results, and the edited results that receive the most support will be distributed. The distribution department can collect viewer feedback in real time and continuously improve the quality of the edited results. Furthermore, the distribution department supports multiple streaming platforms and can distribute the edited results on the platform of the viewer's choice. This allows viewers to watch videos according to their preferences and the platform they are using. Based on viewer feedback, the distribution department can flexibly adjust the content of the distribution and provide the optimal distribution that meets viewer needs. For example, if a viewer wants to add a specific scene, the edited results can be modified and redistributed according to that request. In this way, the distribution department can provide high-quality videos that reflect viewer opinions and improve the viewing experience.
[0034] The extraction unit can extract important scenes from live video footage of sporting events. For example, the extraction unit can automatically detect and extract scoring scenes and highlight scenes. The extraction unit uses image recognition technology and audio analysis technology to detect specific events and actions in the video and extract scenes based on them. For example, the extraction unit can detect goal scenes and decisive plays and provide them to viewers. In this way, by extracting important scenes from live video footage of sporting events, important scenes can be provided to viewers. Some or all of the above processing in the extraction unit may be performed using AI or not. For example, the extraction unit can input live video footage of a sporting event into a generating AI and have the generating AI perform the extraction of important scenes.
[0035] The editorial team allows viewers to select and edit scenes to their liking. For example, viewers can change the order of scenes or add effects. The editorial team provides online collaboration tools and real-time editing functions, enabling viewers to work together on the editing process. For example, viewers can cut and trim scenes and adjust the flow of the video. This increases viewer satisfaction by allowing viewers to select and edit scenes to their liking. Some or all of the above processes in the editorial team may or may not be performed using AI. For example, the editorial team can input scenes selected by viewers into a generating AI and have the generating AI execute editing suggestions.
[0036] The suggestion unit can prioritize suggesting emotionally moving scenes based on the viewer's preferences. For example, the suggestion unit can identify the viewer's preferences based on their past viewing history and survey results, and then make editing suggestions based on those preferences. If the viewer likes emotionally moving scenes, the suggestion unit can prioritize suggesting such scenes. The suggestion unit uses AI to analyze the viewer's preferences and propose the optimal editing. For example, the suggestion unit can prioritize suggesting scenes that the viewer has previously given high ratings to. This increases viewer satisfaction by prioritizing emotionally moving scenes based on the viewer's preferences. Some or all of the above processing in the suggestion unit may be performed using AI or not. For example, the suggestion unit can input the viewer's viewing history data into a generating AI and have the generating AI suggest emotionally moving scenes.
[0037] The distribution department can deliver the edited results to a streaming platform after obtaining viewer approval. The distribution department makes approvals based on viewer votes and feedback and delivers the edited results. The distribution department enables viewers to share videos they have edited with other viewers. For example, the distribution department can upload videos edited by viewers to a streaming platform so that other viewers can watch them. In this way, by delivering the edited results to the streaming platform after obtaining viewer approval, viewers can share videos they have edited. Some or all of the above processes in the distribution department may be performed using AI or not. For example, the distribution department can input viewer feedback data into a generating AI and have the generating AI perform the approval for distribution.
[0038] The Revenue Division can share revenue through Pro Plan subscriptions and advertising integration. For example, through Pro Plan subscriptions, viewers can use more advanced editing functions. Through advertising integration, the Revenue Division can earn revenue by inserting advertisements into videos edited by viewers. For example, the Revenue Division can automatically insert advertisements related to videos edited by viewers and distribute the advertising revenue. This enables a sustainable business model by sharing revenue through Pro Plan subscriptions and advertising integration. Some or all of the above processes in the Revenue Division may be performed using AI or not. For example, the Revenue Division can input viewer editing data into a generating AI and have the generating AI insert advertisements.
[0039] The extraction unit can analyze past viewing data and prioritize extracting scenes that viewers are particularly interested in. For example, the extraction unit can prioritize extracting scenes that viewers have watched multiple times in the past. The extraction unit can prioritize extracting scenes that viewers have given high ratings to in the past. The extraction unit can prioritize extracting scenes that viewers have left many comments on in the past. In this way, by analyzing past viewing data, scenes that viewers are particularly interested in can be prioritized. Some or all of the above processing in the extraction unit may be performed using AI or not. For example, the extraction unit can input the viewer's past viewing data into a generating AI and have the generating AI perform the extraction of scenes of high interest.
[0040] The extraction unit can perform audio analysis of video and extract scenes containing specific keywords. For example, the extraction unit can extract scenes containing keywords such as "goal" or "score" in a sports event. It can extract scenes containing keywords such as "chorus" or "chorus" in a music concert. It can extract scenes containing keywords such as "climax" or "ending" in a movie. In this way, by performing audio analysis of video, scenes containing specific keywords can be extracted. Some or all of the above processing in the extraction unit may be performed using AI or not. For example, the extraction unit can input the audio data of the video into a generating AI and have the generating AI perform the extraction of scenes containing specific keywords.
[0041] The extraction unit can prioritize the extraction of region-specific scenes by taking into account the viewer's geographical location information. For example, if the viewer is in Japan, the extraction unit can prioritize the extraction of scenes related to Japanese scenery and culture. If the viewer is in the United States, the extraction unit can prioritize the extraction of scenes related to American scenery and culture. If the viewer is in Europe, the extraction unit can prioritize the extraction of scenes related to European scenery and culture. In this way, region-specific scenes can be prioritized by taking into account the viewer's geographical location information. Some or all of the above processing in the extraction unit may be performed using AI or not. For example, the extraction unit can input the viewer's geographical location information into a generating AI and have the generating AI perform the extraction of region-specific scenes.
[0042] The extraction unit can analyze the viewer's social media activity and extract relevant scenes. For example, if the viewer frequently posts about "sports" on social media, the extraction unit can prioritize extracting sports scenes. If the viewer frequently posts about "music" on social media, the extraction unit can prioritize extracting music scenes. If the viewer frequently posts about "movies" on social media, the extraction unit can prioritize extracting movie scenes. In this way, relevant scenes can be extracted by analyzing the viewer's social media activity. Some or all of the above processing in the extraction unit may be performed using AI or not. For example, the extraction unit can input the viewer's social media data into a generating AI and have the generating AI perform the extraction of relevant scenes.
[0043] The editorial team can suggest the optimal editing method by referring to the viewer's past editing history. For example, the editorial team can prioritize suggesting effects and music that the viewer has used in the past. The editorial team can suggest the optimal editing method based on the editing style that the viewer has preferred in the past. The editorial team can prioritize suggesting editing methods that the viewer has given high ratings to in the past. In this way, the editorial team can suggest the optimal editing method by referring to the viewer's past editing history. Some or all of the above processes in the editorial team may be performed using AI or not. For example, the editorial team can input the viewer's past editing history data into a generating AI and have the generating AI perform the task of suggesting the optimal editing method.
[0044] The editorial team can filter the content based on the viewer's current interests. For example, if the viewer is currently interested in "sports," the editorial team will prioritize editing sports-related scenes. If the viewer is currently interested in "music," the editorial team can prioritize editing music-related scenes. If the viewer is currently interested in "movies," the editorial team can prioritize editing movies-related scenes. By filtering the content based on the viewer's current interests, the editorial team can provide more appropriate content for the viewer. Some or all of the above processing by the editorial team may be performed using AI or not. For example, the editorial team can input viewer interest data into a generating AI and have the generating AI perform the filtering of the content.
[0045] The editorial team can select the optimal editing method by considering the viewer's device information. For example, if the viewer is using a smartphone, the editorial team can provide an editing method that is adapted to the screen size. If the viewer is using a tablet, the editorial team can provide an editing method optimized for a larger screen. If the viewer is using a smartwatch, the editorial team can provide a concise and highly visible editing method. In this way, the optimal editing method can be provided by considering the viewer's device information. Some or all of the above processing by the editorial team may be performed using AI or not. For example, the editorial team can input the viewer's device information into a generating AI and have the generating AI select the optimal editing method.
[0046] The editorial team can analyze viewers' social media activity and suggest relevant editing methods. For example, if a viewer posts frequently about "sports" on social media, the editorial team can suggest sports-related editing methods. If a viewer posts frequently about "music" on social media, the editorial team can suggest music-related editing methods. If a viewer posts frequently about "movies" on social media, the editorial team can suggest movie-related editing methods. In this way, relevant editing methods can be suggested by analyzing viewers' social media activity. Some or all of the above processing by the editorial team may be performed using AI or not. For example, the editorial team can input viewers' social media data into a generating AI and have the generating AI suggest relevant editing methods.
[0047] The suggestion unit can make optimal suggestions by referring to the viewer's past viewing history. For example, the suggestion unit can prioritize suggesting scenes that the viewer has previously given high ratings to. The suggestion unit can prioritize suggesting scenes that the viewer has replayed many times in the past. The suggestion unit can prioritize suggesting scenes that the viewer has previously commented on many times. In this way, the suggestion unit can make optimal suggestions by referring to the viewer's past viewing history. Some or all of the above processing in the suggestion unit may be performed using AI or not. For example, the suggestion unit can input the viewer's past viewing history data into a generating AI and have the generating AI execute the optimal suggestion.
[0048] The suggestion unit can filter suggestions based on the viewer's current areas of interest. For example, if the viewer is currently interested in "sports," the suggestion unit will prioritize suggesting sports-related scenes. If the viewer is currently interested in "music," the suggestion unit can prioritize suggesting music-related scenes. If the viewer is currently interested in "movies," the suggestion unit can prioritize suggesting movie-related scenes. By filtering suggestions based on the viewer's current areas of interest, the suggestion unit can provide more appropriate suggestions to the viewer. Some or all of the above processing in the suggestion unit may be performed using AI or not. For example, the suggestion unit can input viewer area of interest data into a generating AI and have the generating AI perform the suggestion filtering.
[0049] The suggestion unit can select the optimal suggestion method by considering the viewer's device information. For example, if the viewer is using a smartphone, the suggestion unit can provide a suggestion method that matches the screen size. If the viewer is using a tablet, the suggestion unit can provide a suggestion method optimized for a larger screen. If the viewer is using a smartwatch, the suggestion unit can provide a concise and highly visible suggestion method. In this way, the optimal suggestion method can be provided by considering the viewer's device information. Some or all of the above processing in the suggestion unit may be performed using AI or not. For example, the suggestion unit can input the viewer's device information into a generating AI and have the generating AI select the optimal suggestion method.
[0050] The suggestion unit can analyze viewers' social media activity and make relevant suggestions. For example, if a viewer frequently posts about "sports" on social media, the suggestion unit can prioritize suggesting sports-related scenes. If a viewer frequently posts about "music" on social media, the suggestion unit can prioritize suggesting music-related scenes. If a viewer frequently posts about "movies" on social media, the suggestion unit can prioritize suggesting movie-related scenes. In this way, relevant suggestions can be made by analyzing viewers' social media activity. Some or all of the above processing in the suggestion unit may be performed using AI or not. For example, the suggestion unit can input viewers' social media data into a generating AI and have the generating AI execute relevant suggestions.
[0051] The distribution unit can select the optimal distribution method by referring to the viewer's past viewing history. For example, the distribution unit can prioritize distribution methods that viewers have previously given high ratings to. The distribution unit can prioritize distribution methods that viewers have watched multiple times in the past. The distribution unit can prioritize distribution methods that viewers have left many comments on in the past. In this way, the optimal distribution method can be selected by referring to the viewer's past viewing history. Some or all of the above processing in the distribution unit may be performed using AI or not. For example, the distribution unit can input the viewer's past viewing history data into a generating AI and have the generating AI perform the selection of the optimal distribution method.
[0052] The distribution unit can filter content based on the viewer's current areas of interest. For example, if a viewer is currently interested in "sports," the distribution unit will prioritize sports-related content. If a viewer is currently interested in "music," the distribution unit will prioritize music-related content. If a viewer is currently interested in "movies," the distribution unit will prioritize movie-related content. By filtering content based on the viewer's current areas of interest, the distribution unit can deliver content that is more relevant to the viewer. Some or all of the above processing in the distribution unit may be performed using AI, or not. For example, the distribution unit can input viewer area of interest data into a generating AI and have the generating AI perform the content filtering.
[0053] The distribution unit can select the optimal distribution method by considering the viewer's device information. For example, if the viewer is using a smartphone, the distribution unit can provide a distribution method that matches the screen size. If the viewer is using a tablet, the distribution unit can provide a distribution method optimized for a larger screen. If the viewer is using a smartwatch, the distribution unit can provide a concise and highly visible distribution method. In this way, the optimal distribution method can be provided by considering the viewer's device information. Some or all of the above processing in the distribution unit may be performed using AI or not. For example, the distribution unit can input the viewer's device information into a generating AI and have the generating AI select the optimal distribution method.
[0054] The distribution department can analyze viewers' social media activity and suggest relevant distribution methods. For example, if a viewer frequently posts about "sports" on social media, the distribution department can prioritize sports-related distribution. If a viewer frequently posts about "music" on social media, the distribution department can prioritize music-related distribution. If a viewer frequently posts about "movies" on social media, the distribution department can prioritize movie-related distribution. In this way, by analyzing viewers' social media activity, relevant distribution methods can be suggested. Some or all of the above processing in the distribution department may be performed using AI or not. For example, the distribution department can input viewers' social media data into a generating AI and have the generating AI suggest relevant distribution methods.
[0055] The revenue generation unit can select the optimal revenue method by referring to the viewer's past purchase history. For example, the revenue generation unit can display advertisements related to products the viewer has purchased in the past. The revenue generation unit can display advertisements related to products the viewer has given high ratings to in the past. The revenue generation unit can display advertisements related to products the viewer has purchased multiple times in the past. In this way, the optimal revenue method can be selected by referring to the viewer's past purchase history. Some or all of the above processing in the revenue generation unit may be performed using AI or not. For example, the revenue generation unit can input the viewer's past purchase history data into a generating AI and have the generating AI select the optimal revenue method.
[0056] The revenue generation unit can select the optimal revenue method by considering the viewer's geographical location. For example, if the viewer is in Japan, the revenue generation unit can display advertisements related to Japanese companies and products. If the viewer is in the United States, the revenue generation unit can display advertisements related to American companies and products. If the viewer is in Europe, the revenue generation unit can display advertisements related to European companies and products. In this way, the optimal revenue method can be provided by considering the viewer's geographical location. Some or all of the above processing in the revenue generation unit may be performed using AI or not. For example, the revenue generation unit can input the viewer's geographical location information into a generating AI and have the generating AI select the optimal revenue method.
[0057] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0058] By analyzing a viewer's past viewing history, it is possible to automatically generate video thumbnails based on themes that the viewer is particularly interested in. For example, if a viewer has watched a lot of sports-related videos in the past, thumbnails emphasizing sports scenes can be generated. If a viewer has watched a lot of music-related videos, thumbnails emphasizing music scenes can be generated. If a viewer has watched a lot of movie-related videos, thumbnails emphasizing movie scenes can be generated. In this way, by analyzing a viewer's past viewing history, thumbnails based on themes that the viewer is particularly interested in can be automatically generated. Some or all of the above processes in thumbnail generation may be performed using AI or not. For example, thumbnail generation can be performed by inputting the viewer's past viewing history data into a generation AI and having the generation AI perform the thumbnail generation.
[0059] By considering the viewer's geographical location, region-specific advertisements can be inserted. For example, if the viewer is in Japan, advertisements related to Japanese companies and products can be inserted. If the viewer is in the United States, advertisements related to American companies and products can be inserted. If the viewer is in Europe, advertisements related to European companies and products can be inserted. In this way, region-specific advertisements can be inserted by considering the viewer's geographical location. Some or all of the above processes in ad insertion may be performed using AI or not. For example, ad insertion can be performed by inputting the viewer's geographical location information into a generating AI and having the generating AI perform the ad insertion.
[0060] By analyzing viewers' social media activity, it is possible to automatically generate tags for relevant videos. For example, if a viewer frequently posts about "sports" on social media, sports-related tags can be generated. If a viewer frequently posts about "music" on social media, music-related tags can be generated. If a viewer frequently posts about "movies" on social media, movie-related tags can be generated. In this way, by analyzing viewers' social media activity, tags for relevant videos can be automatically generated. Some or all of the above-described processes in tag generation may be performed using AI or not. For example, tag generation can be performed by inputting the viewer's social media data into a generation AI and having the generation AI execute the tag generation.
[0061] By referring to a viewer's past purchase history, products that are likely to interest the viewer can be introduced in the video. For example, if a viewer has purchased many sporting goods in the past, sporting goods can be introduced. If a viewer has purchased many music-related products, music-related products can be introduced. If a viewer has purchased many movie-related products, movie-related products can be introduced. In this way, by referring to a viewer's past purchase history, products that are likely to interest the viewer can be introduced in the video. Some or all of the above processing in product introduction may be performed using AI or not. For example, product introduction can be performed by inputting the viewer's past purchase history data into a generating AI and having the generating AI execute the product introduction.
[0062] The optimal video format can be selected by considering the viewer's device information. For example, if the viewer is using a smartphone, a video format that matches the screen size can be provided. If the viewer is using a tablet, a video format optimized for a larger screen can be provided. If the viewer is using a smartwatch, a concise and highly visible video format can be provided. In this way, the optimal video format can be provided by considering the viewer's device information. Some or all of the above processes in video format selection may be performed using AI or not. For example, in video format selection, the viewer's device information can be input into a generating AI, and the generating AI can be made to select the optimal video format.
[0063] The following briefly describes the processing flow for example form 1.
[0064] Step 1: The extraction unit automatically extracts multiple scenes from live events and streamed video. The extraction unit uses image recognition and audio analysis technologies to detect specific events and actions within the video and extracts scenes based on them. For example, it can automatically detect and extract scoring scenes and highlight scenes from live video of a sports event. Step 2: The editorial team allows viewers to collaboratively edit the scenes extracted by the extraction team. The editorial team provides online collaboration tools and real-time editing functions, enabling viewers to work together on the editing process. For example, viewers can change the order of scenes or add effects. Step 3: The proposal team dynamically edits the video edited by the editorial team to match the viewer's preferences. The proposal team identifies viewer preferences based on their past viewing history and survey results, and makes editing suggestions accordingly. The proposal team uses AI to analyze viewer preferences and propose the optimal editing. Step 4: The distribution team approves the editing results proposed by the proposal team and distributes them to the streaming platform. By distributing the editing results after viewer approval, the distribution team enables viewers to share the video they helped edit with other viewers. The distribution team approves and distributes the editing results based on viewer votes and feedback.
[0065] (Example of form 2) The automated video editing platform according to an embodiment of the present invention is a system that utilizes generative AI and advanced video analysis to automatically extract multiple scenes from live events and streamed video, and allows viewers to collaboratively edit them. This system uses AI to analyze trends and sentiment, and proposes dynamic editing tailored to each viewer's preferences. The collaborative results are delivered to the streaming platform after approval. The main revenue sources are subscriptions to the Pro plan and revenue sharing through advertising partnerships. For example, the AI automatically extracts multiple scenes from live events and streamed video. For instance, it can extract important scenes from live footage of a sports event. Next, viewers can collaboratively edit the video. In this process, viewers can select and edit scenes according to their preferences. Furthermore, the AI analyzes trends and sentiment, and proposes dynamic editing tailored to each viewer's preferences. For example, if a viewer prefers emotionally moving scenes, the AI can prioritize suggesting such scenes. In this way, viewers can edit according to their preferences. The collaborative results are delivered to the streaming platform after viewer approval. This allows viewers to share the videos they participated in editing with other viewers. Furthermore, the main sources of revenue are subscriptions to the Pro plan and revenue sharing through advertising integration. The Pro plan subscription allows viewers to use more advanced editing features. Additionally, advertising integration allows revenue to be generated by inserting advertisements into videos edited by viewers. In this way, the automated video editing platform, utilizing generative AI and advanced video analysis, enables collaborative editing among viewers and increases viewer satisfaction by suggesting dynamic edits tailored to each viewer's preferences. Moreover, the revenue sharing through Pro plan subscriptions and advertising integration enables a sustainable business model. Thus, the automated video editing platform enables collaborative editing among viewers and suggests dynamic edits tailored to each viewer's preferences.
[0066] The automated video editing platform according to this embodiment comprises an extraction unit, an editing unit, a suggestion unit, and a distribution unit. The extraction unit automatically extracts multiple scenes from live events or streamed video. For example, the extraction unit can extract important scenes from live video of a sports event. The extraction unit uses image recognition technology and audio analysis technology to detect specific events or actions in the video and extracts scenes based on them. For example, the extraction unit can automatically detect and extract scoring scenes or highlight scenes. The editing unit allows viewers to collaboratively edit the scenes extracted by the extraction unit. For example, the editing unit allows viewers to select and edit scenes according to their preferences. The editing unit provides online collaboration tools and real-time editing functions to enable viewers to work together on editing. For example, the editing unit allows viewers to change the order of scenes or add effects. The suggestion unit dynamically edits the video edited by the editing unit according to the viewer's preferences. For example, the suggestion unit identifies the viewer's preferences based on the viewer's past viewing history or survey results and makes editing suggestions based on them. The suggestion unit can prioritize suggesting emotionally impactful scenes if viewers prefer them. The suggestion unit uses AI to analyze viewers' preferences and proposes optimal editing. The distribution unit approves the editing results proposed by the suggestion unit and distributes them to the streaming platform. By distributing the editing results after viewer approval, the distribution unit enables viewers to share the video they participated in editing with other viewers. The distribution unit approves and distributes the editing results based on viewer votes and feedback. As a result, the automated video editing platform according to this embodiment can automatically extract multiple scenes from live events and streamed video, allow viewers to collaboratively edit them, propose dynamic editing tailored to viewers' preferences, and distribute it to the streaming platform.
[0067] The extraction unit automatically extracts multiple scenes from live events and streamed video. Specifically, it can extract important scenes from live video of sports events. The extraction unit uses image recognition and audio analysis technologies to detect specific events and actions within the video and extracts scenes based on that. For example, the extraction unit can automatically detect and extract scoring scenes and highlight scenes. Image recognition technology utilizes object detection algorithms based on deep learning to identify specific movements and objects within the video. For example, in sports events, it can track the movements of players and the position of the ball to identify scoring scenes and important plays. Audio analysis technology analyzes audio signals and detects changes in the tone of cheers and commentary to identify important scenes. For example, it can detect moments when the crowd's cheers get louder or when the commentator is excited and extract those scenes. Furthermore, the extraction unit can integrate video from multiple camera angles and select the optimal scene. This allows for the provision of the most engaging video for viewers. Because the extraction unit analyzes video in real time and instantly extracts important scenes, it can respond quickly even during live events. This allows viewers to enjoy the show in real time without missing any important scenes.
[0068] The editorial team allows viewers to collaboratively edit scenes extracted by the extraction team. Specifically, viewers can select and edit scenes to their liking. The editorial team provides online collaboration tools and real-time editing functions, enabling viewers to work together on the editing process. For example, the editorial team allows viewers to change the order of scenes or add effects. Online collaboration tools enable multiple viewers to work on editing simultaneously and share changes in real time. Viewers can use chat and comment functions to exchange opinions with other viewers as they work on the editing process. The real-time editing function allows viewers to instantly check their edits and make corrections as needed. For example, if a viewer changes the order of scenes, the change is immediately reflected, and the viewer can check the editing results in real time. Furthermore, the editorial team provides a variety of effects and transitions for viewers to use, allowing them to customize the video to their liking. This enables viewers to create their own original videos and share them with other viewers. The editorial team can draw out the creativity of viewers and provide them with the enjoyment of video editing.
[0069] The suggestion department dynamically edits the videos edited by the editorial department to match the viewer's preferences. Specifically, it identifies viewer preferences based on their past viewing history and survey results, and makes editing suggestions accordingly. If a viewer likes emotionally moving scenes, the suggestion department can prioritize suggesting such scenes. The suggestion department uses AI to analyze viewer preferences and propose the optimal editing. The AI analyzes the viewer's viewing history and survey results to identify their preferences and interests. For example, if a viewer has watched many emotionally moving scenes in the past, the AI will determine that the viewer likes emotionally moving scenes and will prioritize suggesting such scenes. The AI can also analyze the viewer's real-time reactions to understand their interests and concerns. For example, if a viewer watches a particular scene for a long time, the AI will determine that the scene is interesting to the viewer and can suggest similar scenes. Furthermore, the suggestion department can collect viewer feedback and continuously improve its suggestions. By providing evaluations and comments on the suggested editing, the suggestion department can more accurately understand viewer preferences and reflect them in future suggestions. This allows the proposal department to provide optimal editing tailored to the viewer's preferences, thereby improving the viewing experience.
[0070] The distribution department approves the editing results proposed by the proposal department and distributes them to streaming platforms. Specifically, by distributing the edited results after viewer approval, viewers can share videos they have edited with other viewers. The distribution department approves and distributes the edited results based on viewer votes and feedback. For example, viewers can vote on the edited results, and the edited results that receive the most support will be distributed. The distribution department can collect viewer feedback in real time and continuously improve the quality of the edited results. Furthermore, the distribution department supports multiple streaming platforms and can distribute the edited results on the platform of the viewer's choice. This allows viewers to watch videos according to their preferences and the platform they are using. Based on viewer feedback, the distribution department can flexibly adjust the content of the distribution and provide the optimal distribution that meets viewer needs. For example, if a viewer wants to add a specific scene, the edited results can be modified and redistributed according to that request. In this way, the distribution department can provide high-quality videos that reflect viewer opinions and improve the viewing experience.
[0071] The extraction unit can extract important scenes from live video footage of sporting events. For example, the extraction unit can automatically detect and extract scoring scenes and highlight scenes. The extraction unit uses image recognition technology and audio analysis technology to detect specific events and actions in the video and extract scenes based on them. For example, the extraction unit can detect goal scenes and decisive plays and provide them to viewers. In this way, by extracting important scenes from live video footage of sporting events, important scenes can be provided to viewers. Some or all of the above processing in the extraction unit may be performed using AI or not. For example, the extraction unit can input live video footage of a sporting event into a generating AI and have the generating AI perform the extraction of important scenes.
[0072] The editorial team allows viewers to select and edit scenes to their liking. For example, viewers can change the order of scenes or add effects. The editorial team provides online collaboration tools and real-time editing functions, enabling viewers to work together on the editing process. For example, viewers can cut and trim scenes and adjust the flow of the video. This increases viewer satisfaction by allowing viewers to select and edit scenes to their liking. Some or all of the above processes in the editorial team may or may not be performed using AI. For example, the editorial team can input scenes selected by viewers into a generating AI and have the generating AI execute editing suggestions.
[0073] The suggestion unit can prioritize suggesting emotionally moving scenes based on the viewer's preferences. For example, the suggestion unit can identify the viewer's preferences based on their past viewing history and survey results, and then make editing suggestions based on those preferences. If the viewer likes emotionally moving scenes, the suggestion unit can prioritize suggesting such scenes. The suggestion unit uses AI to analyze the viewer's preferences and propose the optimal editing. For example, the suggestion unit can prioritize suggesting scenes that the viewer has previously given high ratings to. This increases viewer satisfaction by prioritizing emotionally moving scenes based on the viewer's preferences. Some or all of the above processing in the suggestion unit may be performed using AI or not. For example, the suggestion unit can input the viewer's viewing history data into a generating AI and have the generating AI suggest emotionally moving scenes.
[0074] The distribution department can deliver the edited results to a streaming platform after obtaining viewer approval. The distribution department makes approvals based on viewer votes and feedback and delivers the edited results. The distribution department enables viewers to share videos they have edited with other viewers. For example, the distribution department can upload videos edited by viewers to a streaming platform so that other viewers can watch them. In this way, by delivering the edited results to the streaming platform after obtaining viewer approval, viewers can share videos they have edited. Some or all of the above processes in the distribution department may be performed using AI or not. For example, the distribution department can input viewer feedback data into a generating AI and have the generating AI perform the approval for distribution.
[0075] The Revenue Division can share revenue through Pro Plan subscriptions and advertising integration. For example, through Pro Plan subscriptions, viewers can use more advanced editing functions. Through advertising integration, the Revenue Division can earn revenue by inserting advertisements into videos edited by viewers. For example, the Revenue Division can automatically insert advertisements related to videos edited by viewers and distribute the advertising revenue. This enables a sustainable business model by sharing revenue through Pro Plan subscriptions and advertising integration. Some or all of the above processes in the Revenue Division may be performed using AI or not. For example, the Revenue Division can input viewer editing data into a generating AI and have the generating AI insert advertisements.
[0076] The extraction unit can estimate the viewer's emotions and adjust the criteria for extracting important scenes based on the estimated viewer emotions. For example, if the viewer is excited, the extraction unit can prioritize extracting action scenes. If the viewer is moved, the extraction unit can prioritize extracting emotional scenes. If the viewer is relaxed, the extraction unit can prioritize extracting calm scenes. This allows for the provision of more appropriate scenes to the viewer by adjusting the criteria for extracting important scenes based on the viewer's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the extraction unit may be performed using AI or not. For example, the extraction unit can input viewer emotion data into a generative AI and have the generative AI adjust the criteria for extracting important scenes.
[0077] The extraction unit can analyze past viewing data and prioritize extracting scenes that viewers are particularly interested in. For example, the extraction unit can prioritize extracting scenes that viewers have watched multiple times in the past. The extraction unit can prioritize extracting scenes that viewers have given high ratings to in the past. The extraction unit can prioritize extracting scenes that viewers have left many comments on in the past. In this way, by analyzing past viewing data, scenes that viewers are particularly interested in can be prioritized. Some or all of the above processing in the extraction unit may be performed using AI or not. For example, the extraction unit can input the viewer's past viewing data into a generating AI and have the generating AI perform the extraction of scenes of high interest.
[0078] The extraction unit can perform audio analysis of video and extract scenes containing specific keywords. For example, the extraction unit can extract scenes containing keywords such as "goal" or "score" in a sports event. It can extract scenes containing keywords such as "chorus" or "chorus" in a music concert. It can extract scenes containing keywords such as "climax" or "ending" in a movie. In this way, by performing audio analysis of video, scenes containing specific keywords can be extracted. Some or all of the above processing in the extraction unit may be performed using AI or not. For example, the extraction unit can input the audio data of the video into a generating AI and have the generating AI perform the extraction of scenes containing specific keywords.
[0079] The extraction unit can estimate the viewer's emotions and determine the priority of scenes to extract based on the estimated viewer emotions. For example, if the viewer is excited, the extraction unit can prioritize extracting action scenes. If the viewer is moved, the extraction unit can prioritize extracting emotional scenes. If the viewer is relaxed, the extraction unit can prioritize extracting calm scenes. By prioritizing the scenes to extract based on the viewer's emotions, the system can provide the viewer with more appropriate scenes. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the extraction unit may be performed using AI or not. For example, the extraction unit can input viewer emotion data into a generative AI and have the generative AI determine the priority of scenes.
[0080] The extraction unit can prioritize the extraction of region-specific scenes by taking into account the viewer's geographical location information. For example, if the viewer is in Japan, the extraction unit can prioritize the extraction of scenes related to Japanese scenery and culture. If the viewer is in the United States, the extraction unit can prioritize the extraction of scenes related to American scenery and culture. If the viewer is in Europe, the extraction unit can prioritize the extraction of scenes related to European scenery and culture. In this way, region-specific scenes can be prioritized by taking into account the viewer's geographical location information. Some or all of the above processing in the extraction unit may be performed using AI or not. For example, the extraction unit can input the viewer's geographical location information into a generating AI and have the generating AI perform the extraction of region-specific scenes.
[0081] The extraction unit can analyze the viewer's social media activity and extract relevant scenes. For example, if the viewer frequently posts about "sports" on social media, the extraction unit can prioritize extracting sports scenes. If the viewer frequently posts about "music" on social media, the extraction unit can prioritize extracting music scenes. If the viewer frequently posts about "movies" on social media, the extraction unit can prioritize extracting movie scenes. In this way, relevant scenes can be extracted by analyzing the viewer's social media activity. Some or all of the above processing in the extraction unit may be performed using AI or not. For example, the extraction unit can input the viewer's social media data into a generating AI and have the generating AI perform the extraction of relevant scenes.
[0082] The editorial team can estimate the viewer's emotions and adjust the editing style based on those estimated emotions. For example, if the viewer is moved, the editorial team can add emotional music or effects. If the viewer is excited, the editorial team can add fast-paced music or effects. If the viewer is relaxed, the editorial team can add calming music or effects. By adjusting the editing style based on the viewer's emotions, the editorial team can provide more appropriate editing for the viewer. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the editorial team may be performed using AI or not. For example, the editorial team can input viewer emotion data into a generative AI and have the generative AI adjust the editing style.
[0083] The editorial team can suggest the optimal editing method by referring to the viewer's past editing history. For example, the editorial team can prioritize suggesting effects and music that the viewer has used in the past. The editorial team can suggest the optimal editing method based on the editing style that the viewer has preferred in the past. The editorial team can prioritize suggesting editing methods that the viewer has given high ratings to in the past. In this way, the editorial team can suggest the optimal editing method by referring to the viewer's past editing history. Some or all of the above processes in the editorial team may be performed using AI or not. For example, the editorial team can input the viewer's past editing history data into a generating AI and have the generating AI perform the task of suggesting the optimal editing method.
[0084] The editorial team can filter the content based on the viewer's current interests. For example, if the viewer is currently interested in "sports," the editorial team will prioritize editing sports-related scenes. If the viewer is currently interested in "music," the editorial team can prioritize editing music-related scenes. If the viewer is currently interested in "movies," the editorial team can prioritize editing movies-related scenes. By filtering the content based on the viewer's current interests, the editorial team can provide more appropriate content for the viewer. Some or all of the above processing by the editorial team may be performed using AI or not. For example, the editorial team can input viewer interest data into a generating AI and have the generating AI perform the filtering of the content.
[0085] The editorial team can estimate the viewer's emotions and adjust the order of scenes to be edited based on those estimated emotions. For example, if the viewer is moved, the editorial team can place emotional scenes first. If the viewer is excited, the editorial team can place action scenes first. If the viewer is relaxed, the editorial team can place calm scenes first. By adjusting the order of scenes to be edited based on the viewer's emotions, the editorial team can provide a more appropriate edit for the viewer. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the editorial team may be performed using AI or not. For example, the editorial team can input viewer emotion data into a generative AI and have the generative AI adjust the order of scenes.
[0086] The editorial team can select the optimal editing method by considering the viewer's device information. For example, if the viewer is using a smartphone, the editorial team can provide an editing method that is adapted to the screen size. If the viewer is using a tablet, the editorial team can provide an editing method optimized for a larger screen. If the viewer is using a smartwatch, the editorial team can provide a concise and highly visible editing method. In this way, the optimal editing method can be provided by considering the viewer's device information. Some or all of the above processing by the editorial team may be performed using AI or not. For example, the editorial team can input the viewer's device information into a generating AI and have the generating AI select the optimal editing method.
[0087] The editorial team can analyze viewers' social media activity and suggest relevant editing methods. For example, if a viewer posts frequently about "sports" on social media, the editorial team can suggest sports-related editing methods. If a viewer posts frequently about "music" on social media, the editorial team can suggest music-related editing methods. If a viewer posts frequently about "movies" on social media, the editorial team can suggest movie-related editing methods. In this way, relevant editing methods can be suggested by analyzing viewers' social media activity. Some or all of the above processing by the editorial team may be performed using AI or not. For example, the editorial team can input viewers' social media data into a generating AI and have the generating AI suggest relevant editing methods.
[0088] The suggestion unit can estimate the viewer's emotions and adjust the way it presents suggestions based on those emotions. For example, if the viewer is emotional, the suggestion unit can emphasize emotional scenes in its suggestions. If the viewer is excited, the suggestion unit can emphasize action scenes in its suggestions. If the viewer is relaxed, the suggestion unit can emphasize calm scenes in its suggestions. By adjusting the way it presents suggestions based on the viewer's emotions, it can provide more appropriate suggestions to the viewer. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the processing described above in the suggestion unit may be performed using AI or not. For example, the suggestion unit can input viewer emotion data into a generative AI and have the generative AI adjust the way it presents suggestions.
[0089] The suggestion unit can make optimal suggestions by referring to the viewer's past viewing history. For example, the suggestion unit can prioritize suggesting scenes that the viewer has previously given high ratings to. The suggestion unit can prioritize suggesting scenes that the viewer has replayed many times in the past. The suggestion unit can prioritize suggesting scenes that the viewer has previously commented on many times. In this way, the suggestion unit can make optimal suggestions by referring to the viewer's past viewing history. Some or all of the above processing in the suggestion unit may be performed using AI or not. For example, the suggestion unit can input the viewer's past viewing history data into a generating AI and have the generating AI execute the optimal suggestion.
[0090] The suggestion unit can filter suggestions based on the viewer's current areas of interest. For example, if the viewer is currently interested in "sports," the suggestion unit will prioritize suggesting sports-related scenes. If the viewer is currently interested in "music," the suggestion unit can prioritize suggesting music-related scenes. If the viewer is currently interested in "movies," the suggestion unit can prioritize suggesting movie-related scenes. By filtering suggestions based on the viewer's current areas of interest, the suggestion unit can provide more appropriate suggestions to the viewer. Some or all of the above processing in the suggestion unit may be performed using AI or not. For example, the suggestion unit can input viewer area of interest data into a generating AI and have the generating AI perform the suggestion filtering.
[0091] The suggestion unit can estimate the viewer's emotions and adjust the order of suggested scenes based on those emotions. For example, if the viewer is emotional, the suggestion unit can place emotional scenes first. If the viewer is excited, the suggestion unit can place action scenes first. If the viewer is relaxed, the suggestion unit can place calm scenes first. By adjusting the order of suggested scenes based on the viewer's emotions, the suggestion unit can provide more appropriate suggestions to the viewer. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the processing described above in the suggestion unit may be performed using AI or not. For example, the suggestion unit can input viewer emotion data into a generative AI and have the generative AI adjust the order of scenes.
[0092] The suggestion unit can select the optimal suggestion method by considering the viewer's device information. For example, if the viewer is using a smartphone, the suggestion unit can provide a suggestion method that matches the screen size. If the viewer is using a tablet, the suggestion unit can provide a suggestion method optimized for a larger screen. If the viewer is using a smartwatch, the suggestion unit can provide a concise and highly visible suggestion method. In this way, the optimal suggestion method can be provided by considering the viewer's device information. Some or all of the above processing in the suggestion unit may be performed using AI or not. For example, the suggestion unit can input the viewer's device information into a generating AI and have the generating AI select the optimal suggestion method.
[0093] The suggestion unit can analyze viewers' social media activity and make relevant suggestions. For example, if a viewer frequently posts about "sports" on social media, the suggestion unit can prioritize suggesting sports-related scenes. If a viewer frequently posts about "music" on social media, the suggestion unit can prioritize suggesting music-related scenes. If a viewer frequently posts about "movies" on social media, the suggestion unit can prioritize suggesting movie-related scenes. In this way, relevant suggestions can be made by analyzing viewers' social media activity. Some or all of the above processing in the suggestion unit may be performed using AI or not. For example, the suggestion unit can input viewers' social media data into a generating AI and have the generating AI execute relevant suggestions.
[0094] The distribution unit can estimate the emotions of viewers and adjust the timing of distribution based on the estimated emotions. For example, if viewers are excited, the distribution unit can distribute in real time. If viewers are relaxed, the distribution unit can distribute at a time convenient for the viewers. If viewers are moved, the distribution unit can prioritize distribution of emotionally moving scenes. In this way, by adjusting the timing of distribution based on viewers' emotions, distribution can be delivered at a more appropriate time for viewers. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the distribution unit may be performed using AI or not. For example, the distribution unit can input viewer emotion data into a generative AI and have the generative AI adjust the timing of distribution.
[0095] The distribution unit can select the optimal distribution method by referring to the viewer's past viewing history. For example, the distribution unit can prioritize distribution methods that viewers have previously given high ratings to. The distribution unit can prioritize distribution methods that viewers have watched multiple times in the past. The distribution unit can prioritize distribution methods that viewers have left many comments on in the past. In this way, the optimal distribution method can be selected by referring to the viewer's past viewing history. Some or all of the above processing in the distribution unit may be performed using AI or not. For example, the distribution unit can input the viewer's past viewing history data into a generating AI and have the generating AI perform the selection of the optimal distribution method.
[0096] The distribution unit can filter content based on the viewer's current areas of interest. For example, if a viewer is currently interested in "sports," the distribution unit will prioritize sports-related content. If a viewer is currently interested in "music," the distribution unit will prioritize music-related content. If a viewer is currently interested in "movies," the distribution unit will prioritize movie-related content. By filtering content based on the viewer's current areas of interest, the distribution unit can deliver content that is more relevant to the viewer. Some or all of the above processing in the distribution unit may be performed using AI, or not. For example, the distribution unit can input viewer area of interest data into a generating AI and have the generating AI perform the content filtering.
[0097] The distribution unit can estimate the viewer's emotions and adjust the order of scenes to be distributed based on the estimated viewer emotions. For example, if the viewer is moved, the distribution unit can place emotional scenes first. If the viewer is excited, the distribution unit can place action scenes first. If the viewer is relaxed, the distribution unit can place calm scenes first. In this way, by adjusting the order of scenes to be distributed based on the viewer's emotions, a more appropriate distribution can be made for the viewer. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the distribution unit may be performed using AI or not. For example, the distribution unit can input viewer emotion data into a generative AI and have the generative AI perform the adjustment of the scene order.
[0098] The distribution unit can select the optimal distribution method by considering the viewer's device information. For example, if the viewer is using a smartphone, the distribution unit can provide a distribution method that matches the screen size. If the viewer is using a tablet, the distribution unit can provide a distribution method optimized for a larger screen. If the viewer is using a smartwatch, the distribution unit can provide a concise and highly visible distribution method. In this way, the optimal distribution method can be provided by considering the viewer's device information. Some or all of the above processing in the distribution unit may be performed using AI or not. For example, the distribution unit can input the viewer's device information into a generating AI and have the generating AI select the optimal distribution method.
[0099] The distribution department can analyze viewers' social media activity and suggest relevant distribution methods. For example, if a viewer frequently posts about "sports" on social media, the distribution department can prioritize sports-related distribution. If a viewer frequently posts about "music" on social media, the distribution department can prioritize music-related distribution. If a viewer frequently posts about "movies" on social media, the distribution department can prioritize movie-related distribution. In this way, by analyzing viewers' social media activity, relevant distribution methods can be suggested. Some or all of the above processing in the distribution department may be performed using AI or not. For example, the distribution department can input viewers' social media data into a generating AI and have the generating AI suggest relevant distribution methods.
[0100] The revenue unit can estimate the viewer's emotions and adjust the revenue model based on those estimated emotions. For example, if the viewer is emotional, the revenue unit can display ads related to emotional scenes. If the viewer is excited, the revenue unit can display ads related to action scenes. If the viewer is relaxed, the revenue unit can display ads related to calm scenes. This allows for a more appropriate revenue model for the viewer by adjusting the revenue model based on their emotions. Emotion estimation is achieved using emotion estimation functions, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the revenue unit may be performed using AI or not. For example, the revenue unit can input viewer emotion data into a generative AI and have the generative AI perform the adjustment of the revenue model.
[0101] The revenue generation unit can select the optimal revenue method by referring to the viewer's past purchase history. For example, the revenue generation unit can display advertisements related to products the viewer has purchased in the past. The revenue generation unit can display advertisements related to products the viewer has given high ratings to in the past. The revenue generation unit can display advertisements related to products the viewer has purchased multiple times in the past. In this way, the optimal revenue method can be selected by referring to the viewer's past purchase history. Some or all of the above processing in the revenue generation unit may be performed using AI or not. For example, the revenue generation unit can input the viewer's past purchase history data into a generating AI and have the generating AI select the optimal revenue method.
[0102] The revenue unit can estimate the viewer's emotions and prioritize revenue based on those estimated emotions. For example, if the viewer is emotional, the revenue unit may prioritize revenue related to emotionally moving scenes. If the viewer is excited, the revenue unit may prioritize revenue related to action scenes. If the viewer is relaxed, the revenue unit may prioritize revenue related to calm scenes. This allows for a more appropriate revenue approach for the viewer by prioritizing revenue based on their emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the revenue unit may be performed using AI or not. For example, the revenue unit can input viewer emotion data into a generative AI and have the generative AI determine the revenue prioritization.
[0103] The revenue generation unit can select the optimal revenue method by considering the viewer's geographical location. For example, if the viewer is in Japan, the revenue generation unit can display advertisements related to Japanese companies and products. If the viewer is in the United States, the revenue generation unit can display advertisements related to American companies and products. If the viewer is in Europe, the revenue generation unit can display advertisements related to European companies and products. In this way, the optimal revenue method can be provided by considering the viewer's geographical location. Some or all of the above processing in the revenue generation unit may be performed using AI or not. For example, the revenue generation unit can input the viewer's geographical location information into a generating AI and have the generating AI select the optimal revenue method.
[0104] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0105] The system can estimate the viewer's emotions and adjust the video's color tone based on those emotions. For example, if the viewer is moved, warm colors can be emphasized. If the viewer is excited, vibrant colors can be emphasized. If the viewer is relaxed, calm colors can be emphasized. By adjusting the video's color tone based on the viewer's emotions, a more appropriate viewing experience can be provided. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in video color tone adjustment may be performed using AI or not. For example, video color tone adjustment can be performed by inputting viewer emotion data into a generative AI and having the generative AI perform the color tone adjustment.
[0106] By analyzing a viewer's past viewing history, it is possible to automatically generate video thumbnails based on themes that the viewer is particularly interested in. For example, if a viewer has watched a lot of sports-related videos in the past, thumbnails emphasizing sports scenes can be generated. If a viewer has watched a lot of music-related videos, thumbnails emphasizing music scenes can be generated. If a viewer has watched a lot of movie-related videos, thumbnails emphasizing movie scenes can be generated. In this way, by analyzing a viewer's past viewing history, thumbnails based on themes that the viewer is particularly interested in can be automatically generated. Some or all of the above processes in thumbnail generation may be performed using AI or not. For example, thumbnail generation can be performed by inputting the viewer's past viewing history data into a generation AI and having the generation AI perform the thumbnail generation.
[0107] The system can estimate the viewer's emotions and adjust the video volume based on those emotions. For example, if the viewer is moved, the music volume can be increased. If the viewer is excited, the sound effect volume can be increased. If the viewer is relaxed, the overall volume can be decreased. By adjusting the video volume based on the viewer's emotions, a more appropriate audio experience can be provided to the viewer. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in volume adjustment may be performed using AI or not. For example, volume adjustment can be performed by inputting viewer emotion data into a generative AI and having the generative AI perform the volume adjustment.
[0108] By considering the viewer's geographical location, region-specific advertisements can be inserted. For example, if the viewer is in Japan, advertisements related to Japanese companies and products can be inserted. If the viewer is in the United States, advertisements related to American companies and products can be inserted. If the viewer is in Europe, advertisements related to European companies and products can be inserted. In this way, region-specific advertisements can be inserted by considering the viewer's geographical location. Some or all of the above processes in ad insertion may be performed using AI or not. For example, ad insertion can be performed by inputting the viewer's geographical location information into a generating AI and having the generating AI perform the ad insertion.
[0109] By analyzing viewers' social media activity, it is possible to automatically generate tags for relevant videos. For example, if a viewer frequently posts about "sports" on social media, sports-related tags can be generated. If a viewer frequently posts about "music" on social media, music-related tags can be generated. If a viewer frequently posts about "movies" on social media, movie-related tags can be generated. In this way, by analyzing viewers' social media activity, tags for relevant videos can be automatically generated. Some or all of the above-described processes in tag generation may be performed using AI or not. For example, tag generation can be performed by inputting the viewer's social media data into a generation AI and having the generation AI execute the tag generation.
[0110] The system can estimate the viewer's emotions and adjust the video playback speed based on those emotions. For example, if the viewer is moved, the playback speed can be slowed down. If the viewer is excited, the playback speed can be increased. If the viewer is relaxed, the playback speed can be returned to normal. By adjusting the video playback speed based on the viewer's emotions, a more appropriate viewing experience can be provided to the viewer. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in playback speed adjustment may be performed using AI or not. For example, playback speed adjustment can be performed by inputting viewer emotion data into a generative AI and having the generative AI perform the playback speed adjustment.
[0111] By referring to a viewer's past purchase history, products that are likely to interest the viewer can be introduced in the video. For example, if a viewer has purchased many sporting goods in the past, sporting goods can be introduced. If a viewer has purchased many music-related products, music-related products can be introduced. If a viewer has purchased many movie-related products, movie-related products can be introduced. In this way, by referring to a viewer's past purchase history, products that are likely to interest the viewer can be introduced in the video. Some or all of the above processing in product introduction may be performed using AI or not. For example, product introduction can be performed by inputting the viewer's past purchase history data into a generating AI and having the generating AI execute the product introduction.
[0112] The system can estimate the viewer's emotions and adjust the video subtitles based on those emotions. For example, if the viewer is moved, emotional subtitles matching the emotional scene can be displayed. If the viewer is excited, exciting subtitles matching the action scene can be displayed. If the viewer is relaxed, relaxing subtitles matching the calm scene can be displayed. By adjusting the video subtitles based on the viewer's emotions, a more appropriate subtitle experience can be provided to the viewer. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in subtitle adjustment may be performed using AI or not. For example, subtitle adjustment can be performed by inputting viewer emotion data into a generative AI and having the generative AI perform the subtitle adjustment.
[0113] The optimal video format can be selected by considering the viewer's device information. For example, if the viewer is using a smartphone, a video format that matches the screen size can be provided. If the viewer is using a tablet, a video format optimized for a larger screen can be provided. If the viewer is using a smartwatch, a concise and highly visible video format can be provided. In this way, the optimal video format can be provided by considering the viewer's device information. Some or all of the above processes in video format selection may be performed using AI or not. For example, in video format selection, the viewer's device information can be input into a generating AI, and the generating AI can be made to select the optimal video format.
[0114] The system can estimate the viewer's emotions and adjust the video effects based on those emotions. For example, if the viewer is moved, emotional effects can be added. If the viewer is excited, dynamic effects can be added. If the viewer is relaxed, calming effects can be added. By adjusting the video effects based on the viewer's emotions, a more appropriate video experience can be provided to the viewer. Emotion estimation is achieved using emotion estimation functions, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in effect adjustment may be performed using AI or not. For example, effect adjustment can be performed by inputting viewer emotion data into a generative AI and having the generative AI perform the effect adjustment.
[0115] The following briefly describes the processing flow for example form 2.
[0116] Step 1: The extraction unit automatically extracts multiple scenes from live events and streamed video. The extraction unit uses image recognition and audio analysis technologies to detect specific events and actions within the video and extracts scenes based on them. For example, it can automatically detect and extract scoring scenes and highlight scenes from live video of a sports event. Step 2: The editorial team allows viewers to collaboratively edit the scenes extracted by the extraction team. The editorial team provides online collaboration tools and real-time editing functions, enabling viewers to work together on the editing process. For example, viewers can change the order of scenes or add effects. Step 3: The proposal team dynamically edits the video edited by the editorial team to match the viewer's preferences. The proposal team identifies viewer preferences based on their past viewing history and survey results, and makes editing suggestions accordingly. The proposal team uses AI to analyze viewer preferences and propose the optimal editing. Step 4: The distribution team approves the editing results proposed by the proposal team and distributes them to the streaming platform. By distributing the editing results after viewer approval, the distribution team enables viewers to share the video they helped edit with other viewers. The distribution team approves and distributes the editing results based on viewer votes and feedback.
[0117] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0118] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0119] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0120] Each of the multiple elements described above, including the extraction unit, editing unit, proposal unit, distribution unit, and revenue unit, is implemented, for example, by at least one of the smart device 14 and the data processing unit 12. For example, the extraction unit acquires video and audio using the camera 42 and microphone 38B of the smart device 14 and performs video analysis using the specific processing unit 290 of the data processing unit 12. The editing unit allows viewers to select and edit scenes using the control unit 46A of the smart device 14. The proposal unit analyzes viewers' preferences using the specific processing unit 290 of the data processing unit 12 and makes editing suggestions. The distribution unit approves the editing results using the specific processing unit 290 of the data processing unit 12 and distributes them to the streaming platform. The revenue unit shares revenue through pro plan subscriptions and advertising integration using the specific processing unit 290 of the data processing unit 12. The correspondence between each unit and the devices and control units is not limited to the example described above and can be changed in various ways.
[0121] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0122] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0123] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0124] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0125] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0126] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0127] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0128] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0129] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0130] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0131] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0132] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0133] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0134] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0135] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0136] Each of the multiple elements described above, including the extraction unit, editing unit, proposal unit, distribution unit, and revenue unit, is implemented by, for example, at least one of the smart glasses 214 and the data processing unit 12. For example, the extraction unit acquires video and audio using the camera 42 and microphone 238 of the smart glasses 214 and performs video analysis using the specific processing unit 290 of the data processing unit 12. The editing unit allows viewers to select and edit scenes using the control unit 46A of the smart glasses 214. The proposal unit analyzes viewers' preferences using the specific processing unit 290 of the data processing unit 12 and makes editing suggestions. The distribution unit approves the editing results using the specific processing unit 290 of the data processing unit 12 and distributes them to the streaming platform. The revenue unit shares revenue through pro plan subscriptions and advertising integration using the specific processing unit 290 of the data processing unit 12. The correspondence between each unit and the devices and control units is not limited to the example described above and can be changed in various ways.
[0137] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0138] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0139] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0140] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0141] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0142] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0143] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0144] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0145] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0146] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0147] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0148] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0149] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0150] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0151] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0152] Each of the multiple elements described above, including the extraction unit, editing unit, proposal unit, distribution unit, and revenue unit, is implemented by, for example, at least one of the headset terminal 314 and the data processing unit 12. For example, the extraction unit acquires video and audio using the camera 42 and microphone 238 of the headset terminal 314 and performs video analysis using the specific processing unit 290 of the data processing unit 12. The editing unit allows viewers to select and edit scenes using the control unit 46A of the headset terminal 314. The proposal unit analyzes viewers' preferences using the specific processing unit 290 of the data processing unit 12 and makes editing suggestions. The distribution unit approves the editing results using the specific processing unit 290 of the data processing unit 12 and distributes them to the streaming platform. The revenue unit shares revenue through pro plan subscriptions and advertising integration using the specific processing unit 290 of the data processing unit 12. The correspondence between each unit and the devices and control units is not limited to the example described above and can be changed in various ways.
[0153] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0154] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0155] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0156] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0157] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0158] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0159] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0160] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0161] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0162] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0163] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0164] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0165] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0166] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0167] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0168] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0169] Each of the multiple elements described above, including the extraction unit, editing unit, proposal unit, distribution unit, and revenue unit, is implemented by, for example, at least one of the robot 414 and the data processing unit 12. For example, the extraction unit acquires video and audio using the camera 42 and microphone 238 of the robot 414 and performs video analysis using the specific processing unit 290 of the data processing unit 12. The editing unit allows viewers to select and edit scenes using the control unit 46A of the robot 414. The proposal unit analyzes viewers' preferences using the specific processing unit 290 of the data processing unit 12 and makes editing suggestions. The distribution unit approves the editing results using the specific processing unit 290 of the data processing unit 12 and distributes them to the streaming platform. The revenue unit shares revenue through pro plan subscriptions and advertising integration using the specific processing unit 290 of the data processing unit 12. The correspondence between each unit and the devices and control units is not limited to the example described above and can be changed in various ways.
[0170] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0171] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0172] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0173] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0174] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0175] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0176] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0177] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0178] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0179] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0180] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0181] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0182] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0183] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0184] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0185] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0186] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0187] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0188] (Note 1) An extraction unit that automatically extracts multiple scenes from live events and streamed video, The editing department, in which viewers collaboratively edit the scenes extracted by the aforementioned extraction unit, The aforementioned editorial department has a proposal department that dynamically edits the video to suit the viewer's preferences, The distribution unit approves the editing results proposed by the proposal unit and distributes them to a streaming platform. A system characterized by the following features. (Note 2) The extraction unit is Extract key scenes from live footage of sporting events. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned editorial department, Viewers can select and edit scenes to suit their own preferences. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned proposal section is, The program prioritizes suggesting emotionally impactful scenes to match the viewer's preferences. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned distribution unit, The edited version will be distributed to the streaming platform after viewer approval. The system described in Appendix 1, characterized by the features described herein. (Note 6) It has a revenue-generating division that shares revenue through Pro plan subscriptions and advertising partnerships. The system described in Appendix 1, characterized by the features described herein. (Note 7) The extraction unit is We estimate the audience's emotions and adjust the criteria for extracting important scenes based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The extraction unit is By analyzing past viewing data, the system prioritizes extracting scenes that viewers are particularly interested in. The system described in Appendix 1, characterized by the features described herein. (Note 9) The extraction unit is The system analyzes the audio of the video and extracts scenes that contain specific keywords. The system described in Appendix 1, characterized by the features described herein. (Note 10) The extraction unit is It estimates the viewer's emotions and determines the priority of scenes to extract based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The extraction unit is The system prioritizes extracting region-specific scenes by considering the viewer's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 12) The extraction unit is Analyze viewers' social media activity and extract relevant scenes. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned editorial department, The system estimates the audience's emotions and adjusts the editing style based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned editorial department, We suggest the best editing method by referring to the viewer's past editing history. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned editorial department, Filter the editing based on the viewer's current areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned editorial department, It estimates the viewer's emotions and adjusts the order of scenes to be edited based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned editorial department, The optimal editing method is selected considering the viewer's device information. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned editorial department, We analyze viewers' social media activity and suggest relevant editing methods. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned proposal section is, We estimate the audience's emotions and adjust the way we present our proposals based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned proposal section is, We make optimal suggestions by referring to the viewer's past viewing history. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned proposal section is, Filter suggestions based on the viewer's current areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned proposal section is, It estimates the viewer's emotions and adjusts the suggested scene order based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned proposal section is, The optimal suggestion method is selected considering the viewer's device information. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned proposal section is, Analyze viewers' social media activity and make relevant suggestions. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned distribution unit, We estimate the audience's emotions and adjust the timing of the broadcast based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned distribution unit, The optimal delivery method is selected by referring to the viewer's past viewing history. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned distribution unit, Filter content based on the viewer's current areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned distribution unit, It estimates the viewer's emotions and adjusts the order of scenes delivered based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned distribution unit, The optimal delivery method is selected considering the viewer's device information. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned distribution unit, We analyze viewers' social media activity and suggest relevant distribution methods. The system described in Appendix 1, characterized by the features described herein. (Note 31) The aforementioned revenue-generating section is, We estimate the audience's emotions and adjust the revenue model based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 32) The aforementioned revenue-generating section is, The optimal revenue method is selected by referring to the viewer's past purchase history. The system described in Appendix 1, characterized by the features described herein. (Note 33) The aforementioned revenue-generating section is, The system estimates viewer sentiment and determines revenue priorities based on that estimated sentiment. The system described in Appendix 1, characterized by the features described herein. (Note 34) The aforementioned revenue-generating section is, Select the optimal revenue method by considering the geographical location of the viewers. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]
[0189] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. An extraction unit that automatically extracts multiple scenes from live events and streamed video, The editing department, in which viewers collaboratively edit the scenes extracted by the aforementioned extraction unit, The aforementioned editorial department has a proposal department that dynamically edits the video to suit the viewer's preferences, The distribution unit approves the editing results proposed by the proposal unit and distributes them to a streaming platform. A system characterized by the following features.
2. The extraction unit is Extract key scenes from live footage of sporting events. The system according to feature 1.
3. The aforementioned editorial department, Viewers can select and edit scenes to suit their own preferences. The system according to feature 1.
4. The aforementioned proposal section is, The program prioritizes suggesting emotionally impactful scenes to match the viewer's preferences. The system according to feature 1.
5. The aforementioned distribution unit, The edited version will be distributed to the streaming platform after viewer approval. The system according to feature 1.
6. It has a revenue-generating division that shares revenue through Pro plan subscriptions and advertising partnerships. The system according to feature 1.
7. The extraction unit is We estimate the audience's emotions and adjust the criteria for extracting important scenes based on those estimated emotions. The system according to feature 1.
8. The extraction unit is By analyzing past viewing data, the system prioritizes extracting scenes that viewers are particularly interested in. The system according to feature 1.
9. The extraction unit is The system analyzes the audio of the video and extracts scenes that contain specific keywords. The system according to feature 1.
10. The extraction unit is It estimates the viewer's emotions and determines the priority of scenes to extract based on those estimated emotions. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A