system
The system addresses the challenge of real-time highlight extraction and editing from multiple video materials by employing AI to identify and edit highlights, generating immersive videos with optimal scene selection and individual incorporation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies face difficulties in extracting highlights from multiple video materials and editing them in real time.
A system comprising multiple microphone-equipped cameras, an editing unit, a generation unit, an authentication unit, and a reading unit, utilizing AI to instantly identify highlights, perform real-time editing, generate complete videos, and incorporate specific individuals or data from smartphones, enabling immersive video generation.
The system efficiently extracts and edits highlights from video materials in real time, generating high-quality, immersive videos that meet customer demands by using AI to select and edit optimal scenes, prioritize specific individuals, and enhance video quality.
Smart Images

Figure 2026072721000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there was a problem that it was difficult to extract highlights from a plurality of video materials and edit them in real time.
[0005] The system according to the embodiment aims to extract highlights from a plurality of video materials and edit them in real time.
Means for Solving the Problems
[0006] The system according to this embodiment comprises an editing unit, a generation unit, an authentication unit, and a reading unit. The editing unit is equipped with multiple microphone-equipped cameras, and AI instantly identifies highlights and performs real-time editing. The generation unit generates a single video from the video material edited by the editing unit. The authentication unit reads photo data, performs facial recognition, and generates a video by specifying the person to be included in the video. The reading unit reads video data shot with a smartphone or the like and generates an immersive video. [Effects of the Invention]
[0007] The system according to this embodiment can extract highlights from multiple video materials and edit them in real time. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10]This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between a plurality of computers. Examples of communication standards applicable to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The video generation system according to an embodiment of the present invention is a system in which multiple microphone-equipped cameras are arranged, and AI instantly identifies highlights and performs real-time editing to complete a single video. This system can generate videos freely from any video material. Furthermore, by reading photo data and performing facial recognition, it is also possible to specify the people to be included in the video and generate it. In addition, video data shot with smartphones can also be read to generate more immersive videos. For example, multiple microphone-equipped cameras are placed in the venue of a wedding reception, and AI instantly identifies highlights and performs real-time editing to complete a single video by the end of the banquet. It can generate videos freely from any video material, and by reading group photos and performing facial recognition, it is also possible to specify the guests to be included in the video and generate it. In addition, video data shot by attendees can also be read to generate more immersive videos. Since it uses wireless LAN-compatible network cameras, it can be applied to events other than wedding receptions, provided the environment is suitable. This system was developed to address the current situation where securing photographers has become difficult amidst the increasing trend in the number of wedding receptions since the end of the COVID-19 pandemic. Customer demand is increasing, and by using AI to select and edit the best scenes, we can meet customer needs. This allows the video generation system to use multiple microphone-equipped cameras, with AI instantly identifying highlights and editing in real time to produce a complete video.
[0029] The video generation system according to this embodiment comprises an editing unit, a generation unit, an authentication unit, and a reading unit. The editing unit arranges multiple cameras with microphones, and the AI instantly identifies highlights and performs real-time editing. For example, the editing unit collects video material using multiple cameras with microphones, and the AI instantly identifies highlights. The editing unit can also edit video material in real time using the AI. For example, the editing unit has the AI select important scenes from the video material and edit them in real time. The generation unit generates a complete video from the video material edited by the editing unit. For example, the generation unit has the AI generate a complete video based on the video material edited by the editing unit. The generation unit can also combine video material using the AI to generate a complete video. For example, the generation unit has the AI select the optimal scenes from the video material and generate a complete video. The authentication unit reads photo data, performs facial recognition, and generates a video specifying the person to be included in the video. For example, the authentication unit has the AI analyze photo data and perform facial recognition. Furthermore, the authentication unit can use AI to recognize specific individuals from photo data and incorporate them into videos. For example, the authentication unit uses AI to recognize a specific person from photo data and incorporate that person into the video. The reading unit reads video data shot with a smartphone or other device and generates immersive videos. For example, the reading unit uses AI to analyze video data shot with a smartphone and generate immersive videos. The reading unit can also use AI to analyze video data and generate immersive videos. For example, the reading unit uses AI to select important scenes from video data and generate immersive videos. As a result, the video generation system according to this embodiment can use multiple microphone-equipped cameras, have AI instantly grasp the highlights, and edit in real time to generate a single work.
[0030] The editorial team deploys multiple cameras equipped with microphones, and AI instantly identifies highlights and performs real-time editing. Specifically, the editorial team strategically places multiple cameras with microphones at event venues and sports stadiums, with each camera collecting video and audio from a different perspective. These cameras are equipped with wide-angle lenses and zoom capabilities, allowing them to capture wide-area footage in high resolution. The collected video and audio data is transmitted in real time to a central server, where AI analyzes this data. The AI uses speech recognition technology to detect audio events such as cheers and applause, and identifies important scenes from the video. For example, in sports events, goal scenes and player interviews are selected as highlights. Furthermore, the AI analyzes movement and color changes in the video to select visually appealing scenes. This allows the editorial team to instantly extract highlights from a vast amount of video data and edit them in real time. The edited video is immediately distributed to viewers, providing an immersive live experience. In addition, the editorial team utilizes the AI's learning capabilities to optimize the editing algorithm based on past editing data, enabling more accurate highlight selection. This allows the editorial team to achieve efficient and high-quality video editing, enabling them to provide compelling content to viewers.
[0031] The generation unit creates a complete video from video footage edited by the editorial unit. Specifically, the generation unit uses AI to combine video footage based on highlight scenes selected by the editorial unit to create a complete video. The AI considers the flow of the video and storytelling elements, optimizing the order of scenes and transitions. For example, in a sports event highlight video, it arranges important moments of the game in chronological order and appropriately inserts audience reactions and player interviews to create a video that resonates with viewers. Furthermore, the generation unit also performs color correction and audio adjustment to improve the overall quality. The AI automatically adjusts the brightness and contrast of the video and removes audio noise to achieve clear sound quality. The generation unit can also add effects and subtitles to the video. For example, it can insert subtitles displaying player names and scores to make the information easy for viewers to understand. In this way, the generation unit can create a visually and aurally superior video based on the material provided by the editorial unit. Moreover, the generation unit utilizes the AI's learning function to optimize the generation algorithm based on past video data, enabling it to efficiently produce higher quality videos. This allows the production unit to provide viewers with consistent, high-quality video content, thereby increasing their satisfaction.
[0032] The authentication unit reads photo data, performs facial recognition, and generates videos featuring the person to be included. Specifically, the authentication unit uses AI to analyze user-provided photo data and facial recognition technology to identify a specific person. The AI extracts facial feature points and matches them with existing facial data in the database to identify the person. For example, facial photos of event participants or athletes can be registered in advance, and the system can automatically detect scenes in the video in which those individuals appear. The authentication unit can understand which scenes feature a specific person and prioritize editing and generating those scenes. This allows users to easily create video works centered around themselves or specific individuals. Furthermore, the authentication unit can use facial recognition technology to analyze the expressions and movements of people in the video and capture changes in emotion. For example, by emphasizing the expressions of joy or frustration of athletes, it can create videos that move viewers. In addition, the authentication unit can continuously improve the accuracy of facial recognition by utilizing the AI's learning function. This allows the authentication unit to accurately identify specific individuals requested by the user and include them in the video work.
[0033] The loading unit reads video data shot on smartphones and other devices and generates immersive videos. Specifically, the loading unit uses AI to analyze video data shot by the user on their smartphone and selects important scenes from the footage. The AI analyzes the movement and changes in sound in the video, detecting visually appealing scenes and audio events. For example, in travel videos, it selects scenes of beautiful scenery and enjoyable activities to create an immersive video. Furthermore, the loading unit also performs color correction and audio adjustment to improve the overall quality. The AI automatically adjusts the brightness and contrast of the video and removes audio noise to achieve clear sound quality. The loading unit can also add effects and subtitles to the video. For example, it can insert subtitles displaying place names and dates into travel videos, making it easier for viewers to understand the information. In this way, the loading unit can generate visually and aurally superior, immersive videos based on video data shot by the user. Moreover, the loading unit utilizes the AI's learning function to optimize the analysis algorithm based on past video data, enabling it to efficiently generate higher quality videos. This allows the loading unit to provide users with consistent, high-quality video content, thereby increasing their satisfaction.
[0034] The video generation system can utilize Wi-Fi-enabled network cameras. Using Wi-Fi-enabled network cameras increases flexibility in installation locations, enabling use in a variety of environments. For example, Wi-Fi-enabled network cameras can be installed at wedding reception venues. They can also be installed at venues for corporate events, sporting events, concerts, and other similar events. Wi-Fi-enabled network cameras have specific specifications, such as being from a particular manufacturer, model, or communication speed. This flexibility in installation locations and use in diverse environments is a key advantage of using Wi-Fi-enabled network cameras.
[0035] The video generation system can be applied to events other than wedding receptions. Because it can be applied to events other than wedding receptions, it can be used for a wide range of purposes. For example, the video generation system can be applied to various events such as corporate events, sporting events, and concerts. In corporate events, multiple cameras with microphones can be placed, and the AI can instantly identify highlights and edit them in real time to create a single video. In sporting events, multiple cameras with microphones can be placed, and the AI can instantly identify highlights and edit them in real time to create a single video by the end of the match. In concerts, multiple cameras with microphones can be placed, and the AI can instantly identify highlights and edit them in real time to create a single video by the end of the concert. As such, the video generation system can be applied to events other than wedding receptions and can be used for a wide range of purposes.
[0036] The editorial team can collect video footage using multiple cameras equipped with microphones and instantly identify highlights. For example, the editorial team can collect video footage using multiple cameras equipped with microphones. For instance, the editorial team can place multiple cameras equipped with microphones at a wedding reception venue to collect video footage. They can also place multiple cameras equipped with microphones at venues such as corporate events, sporting events, and concerts to collect video footage. The editorial team uses AI to instantly identify highlights from the video footage. For example, the editorial team can use AI to select important scenes from the video footage and instantly identify the highlights. This allows for the collection of video footage using multiple cameras equipped with microphones and the instant identification of highlights.
[0037] The generation unit can generate a complete work from video footage edited by the editorial unit. For example, the generation unit uses AI to generate a complete work based on video footage edited by the editorial unit. For instance, the generation unit uses AI to generate a complete work based on video footage of a wedding reception. The generation unit can also use AI to generate a complete work based on video footage of corporate events, sporting events, concerts, etc. The generation unit uses AI to combine video footage to create a complete work. For example, the generation unit uses AI to select the optimal scenes from the video footage and generate a complete work. In this way, a complete work can be generated from edited video footage.
[0038] The authentication unit can read photo data, perform facial recognition, and generate a video featuring the person to be included. For example, the authentication unit can use AI to analyze photo data and perform facial recognition. For instance, it can read a group photo from a wedding reception, use AI to perform facial recognition, and generate a video featuring the guests to be included. The authentication unit can also read photo data from corporate events, sporting events, concerts, etc., use AI to perform facial recognition, and generate a video featuring the person to be included. The authentication unit uses AI to recognize specific individuals from photo data and incorporate them into the video. For example, the authentication unit uses AI to recognize a specific person from photo data and incorporate that person into the video. This allows for facial recognition using photo data and the inclusion of specific individuals in the video.
[0039] The reading unit can read video data shot on smartphones and other devices and generate immersive videos. For example, the reading unit uses AI to analyze video data shot on a smartphone and generate immersive videos. For instance, the reading unit can read video data shot by wedding reception attendees, analyze it with AI, and generate immersive videos. The reading unit can also read video data shot at corporate events, sporting events, concerts, etc., and analyze it with AI to generate immersive videos. The reading unit uses AI to analyze video data and generate immersive videos. For example, the reading unit's AI selects important scenes from the video data and generates immersive videos. This makes it possible to generate immersive videos using video data shot on smartphones and other devices.
[0040] The editorial team can analyze audio data when collecting video footage and prioritize selecting scenes containing specific keywords as highlights. For example, the editorial team can select scenes containing keywords such as "congratulations" or "I love you" in wedding speeches as highlights. The editorial team can also select scenes containing keywords such as "goal" or "victory" in sporting events as highlights. The editorial team can also select scenes containing keywords such as "encore" or "thank you" in concerts as highlights. By analyzing audio data, the editorial team can prioritize selecting scenes containing specific keywords as highlights. Some or all of the above processing by the editorial team may be performed using AI, for example, or without AI. For example, the editorial department can input audio data into a generative AI and have the AI detect specific keywords.
[0041] The editorial team can detect specific actions or gestures when collecting video footage and select highlights based on them. For example, the editorial team could select a scene of the bride and groom kissing at a wedding as a highlight. The editorial team could also select a scene of an athlete making a victory pose at a sporting event as a highlight. The editorial team could also select a scene of an artist waving to the audience at a concert as a highlight. In this way, by detecting specific actions or gestures, highlights can be selected based on them. Some or all of the above processing in the editorial team may be performed using AI, for example, or not using AI. For example, the editorial team could input video data into a generating AI and have the generating AI perform the detection of specific actions or gestures.
[0042] The editorial team can analyze ambient and background sounds when collecting video footage and select scenes containing specific sounds as highlights. For example, the editorial team can select scenes with a lot of applause and cheering at a wedding as highlights. The editorial team can also select scenes with loud crowd cheering at a sporting event as highlights. The editorial team can also select scenes with audience singing at a concert as highlights. By analyzing ambient and background sounds, the editorial team can select scenes containing specific sounds as highlights. Some or all of the above processing by the editorial team may be performed using AI, for example, or without AI. For example, the editorial team can input audio data into a generating AI and have the generating AI perform the detection of specific sounds.
[0043] The editorial department can analyze the color tone and brightness of video footage during collection and select scenes with specific color tones and brightness levels as highlights. For example, the editorial department can select bright and vibrant scenes as highlights at a wedding. For example, the editorial department can select scenes where the vibrant uniforms stand out at a sporting event as highlights. For example, the editorial department can select scenes with beautiful lighting at a concert as highlights. In this way, by analyzing the color tone and brightness of the video, scenes with specific color tones and brightness levels can be selected as highlights. Some or all of the above processing by the editorial department may be performed using AI, for example, or without AI. For example, the editorial department can input video data into a generating AI and have the generating AI perform the detection of specific color tones and brightness levels.
[0044] The generation unit can construct a natural storyline by considering the temporal flow of the video material during generation. For example, the generation unit can edit wedding video material in chronological order to construct a natural storyline. The generation unit can also edit sports event video material in accordance with the progress of the game to construct a natural storyline. The generation unit can also edit concert video material according to the setlist to construct a natural storyline. By considering the temporal flow of the video material, a natural storyline can be constructed. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input video data into a generation AI and have the generation AI perform editing that considers the temporal flow.
[0045] The generation unit can optimize the synchronization of audio and video in video footage during generation. For example, the generation unit can synchronize the audio and video of speeches in wedding video footage. The generation unit can also synchronize the audio and video of commentary in sporting event video footage. The generation unit can also synchronize the audio and video of performances in concert video footage. By optimizing the synchronization of audio and video in video footage, a more cohesive work can be produced. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input audio and video data into a generation AI and have the generation AI perform the synchronization optimization.
[0046] The generation unit can create visually appealing works by combining different camera angles from video footage during the generation process. For example, the generation unit can combine different angles of the bride and groom from wedding video footage to create a work. The generation unit can also combine different angles of athletes from sporting event video footage to create a work. The generation unit can also combine different angles of artists from concert video footage to create a work. By combining different camera angles, it is possible to create visually appealing works. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input video data into a generation AI and have the generation AI execute combinations of different camera angles.
[0047] The generation unit can generate smooth video by adjusting the different frame rates of video material during generation. For example, the generation unit can generate smooth video by adjusting the different frame rates of wedding video material. The generation unit can also generate smooth video by adjusting the different frame rates of sporting event video material. For example, the generation unit can generate smooth video by adjusting the different frame rates of concert video material. In this way, smooth video can be generated by adjusting the different frame rates. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input video data into a generation AI and have the generation AI perform the frame rate adjustment.
[0048] The authentication unit can apply the optimal facial recognition algorithm during authentication, taking into account the resolution of the photo data. For example, the authentication unit can perform detailed facial recognition using high-resolution photo data. For example, the authentication unit can perform detailed facial recognition using high-resolution photo data of a wedding. The authentication unit can also perform simple facial recognition using low-resolution photo data. For example, the authentication unit can perform simple facial recognition using low-resolution photo data of a sporting event. The authentication unit can also perform balanced facial recognition using medium-resolution photo data. For example, the authentication unit can perform balanced facial recognition using medium-resolution photo data of a concert. This improves authentication accuracy by applying the optimal facial recognition algorithm according to the resolution of the photo data. Some or all of the above processing in the authentication unit may be performed using AI, for example, or without AI. For example, the authentication unit can input photo data into a generating AI and have the generating AI perform facial recognition according to the resolution.
[0049] The authentication unit can perform more accurate facial recognition by combining multiple photographic data during authentication. For example, the authentication unit can perform facial recognition by combining photographic data taken from multiple angles. For example, the authentication unit can perform facial recognition by combining photographic data taken from multiple angles at a wedding. The authentication unit can also perform facial recognition by combining photographic data taken at multiple different times. For example, the authentication unit can perform facial recognition by combining photographic data taken at multiple different times at a sporting event. The authentication unit can also perform facial recognition by combining photographic data taken at multiple different locations. For example, the authentication unit can perform facial recognition by combining photographic data taken at multiple different locations at a concert. By combining multiple photographic data, more accurate facial recognition can be achieved. Some or all of the above processing in the authentication unit may be performed using AI, for example, or without AI. For example, the authentication unit can input multiple photographic data into a generating AI and have the generating AI perform facial recognition.
[0050] The authentication unit can prioritize the use of the most recent photograph during authentication, taking into account the date and time the photograph was taken. For example, the authentication unit can perform facial recognition using the most recent photograph. For example, the authentication unit can perform facial recognition using the most recent photograph from a wedding. The authentication unit can also perform facial recognition using past photographs. For example, the authentication unit can perform facial recognition using past photographs from a sporting event. The authentication unit can also perform facial recognition using photographs taken within a specific period. For example, the authentication unit can perform facial recognition using photographs taken within a specific period from a concert. By considering the date and time the photograph was taken, the authentication unit can prioritize the use of the most recent photograph, thereby improving authentication accuracy. Some or all of the above processing in the authentication unit may be performed using AI, for example, or without AI. For example, the authentication unit can input photographic data into a generating AI and have the generating AI perform facial recognition based on the date and time the photograph was taken.
[0051] The authentication unit analyzes the background information of the photo data during authentication and can prioritize the use of photos that contain a specific background. For example, the authentication unit can perform facial recognition using photo data taken at a specific location. For example, it can perform facial recognition using photo data taken at a specific location at a wedding. The authentication unit can also perform facial recognition using photo data taken at a specific event. For example, it can perform facial recognition using photo data taken at a specific sporting event. The authentication unit can also perform facial recognition using photo data taken at a specific time of day. For example, it can perform facial recognition using photo data taken at a specific time of day at a concert. By analyzing the background information of the photo data, the authentication unit can prioritize the use of photos that contain a specific background, thereby improving authentication accuracy. Some or all of the above processing in the authentication unit may be performed using AI, for example, or without AI. For example, the authentication unit can input photo data into a generating AI and have the generating AI perform facial recognition based on the background information.
[0052] The loading unit can select the optimal loading method considering the resolution of the video data during loading. For example, the loading unit can prioritize loading high-resolution video data to provide detailed footage. For example, the loading unit can prioritize loading high-resolution video data of a wedding to provide detailed footage. The loading unit can also prioritize loading low-resolution video data to enable faster playback. For example, the loading unit can prioritize loading low-resolution video data of a sporting event to enable faster playback. The loading unit can also prioritize loading medium-resolution video data to provide balanced footage. For example, the loading unit can prioritize loading medium-resolution video data of a concert to provide balanced footage. In this way, by selecting the optimal loading method according to the resolution of the video data, higher quality video can be provided. Some or all of the above processing in the loading unit may be performed using AI, for example, or without AI. For example, the loading unit can input video data into a generating AI and have the generating AI perform loading according to the resolution.
[0053] The loading unit can adjust the frame rate of video data during loading to achieve smooth playback. For example, the loading unit can prioritize loading high-frame-rate video data to achieve smooth playback. For example, the loading unit can prioritize loading high-frame-rate video data of a wedding to achieve smooth playback. The loading unit can also prioritize loading low-frame-rate video data to achieve fast playback. For example, the loading unit can prioritize loading low-frame-rate video data of a sporting event to achieve fast playback. The loading unit can also prioritize loading medium-frame-rate video data to achieve balanced playback. For example, the loading unit can prioritize loading medium-frame-rate video data of a concert to achieve balanced playback. In this way, smooth playback can be achieved by adjusting the frame rate of the video data. Some or all of the above processing in the loading unit may be performed using AI, for example, or without AI. For example, the loading unit can input video data into a generating AI and have the generating AI perform the frame rate adjustment.
[0054] The loading unit can prioritize loading the latest video data by considering the shooting date and time of the video data during loading. For example, the loading unit can prioritize loading the latest video data and provide the most recent footage. For example, the loading unit can prioritize loading the latest video data of a wedding and provide the most recent footage. The loading unit can also prioritize loading past video data and provide past footage. For example, the loading unit can prioritize loading past video data of a sporting event and provide past footage. The loading unit can also prioritize loading video data shot within a specific period and provide footage from that period. For example, the loading unit can prioritize loading video data shot within a specific period of a concert and provide footage from that period. This allows the loading unit to prioritize loading the latest video data by considering the shooting date and time of the video data. Some or all of the above processing in the loading unit may be performed using AI, for example, or without AI. For example, the loading unit can input video data into a generating AI and have the generating AI perform loading based on the shooting date and time.
[0055] The loading unit can analyze the audio information of video data during loading and prioritize loading videos that contain specific audio. For example, the loading unit can prioritize loading video data that contains audio with specific keywords. For example, it can prioritize loading video data from weddings that contains audio with keywords such as "congratulations" or "I love you." The loading unit can also prioritize loading video data that contains specific music. For example, it can prioritize loading video data from sporting events that contains specific music. The loading unit can also prioritize loading video data that contains specific ambient sounds. For example, it can prioritize loading video data from concerts that contain specific ambient sounds. In this way, by analyzing the audio information of video data, it is possible to prioritize loading videos that contain specific audio. Some or all of the above processing in the loading unit may be performed using AI, for example, or without AI. For example, the loading unit can input audio data into a generating AI and have the generating AI perform the detection of specific audio.
[0056] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0057] The video generation system can also be equipped with the ability to detect specific weather conditions when collecting video footage and automatically add video effects based on those conditions. For example, during rainy weather, a rain effect can be added to the video to enhance realism. During sunny weather, a sunlight effect can be added to create a bright and refreshing atmosphere. Furthermore, on snowy days, a snow effect can be added to emphasize the winter atmosphere. In this way, by automatically adding video effects based on weather conditions, it is possible to generate more realistic and immersive videos.
[0058] The video generation system can also be equipped with the ability to detect the movements of specific animals during the collection of video footage and edit the footage based on that detection. For example, it can detect scenes featuring pets in wedding footage and select those scenes as highlights. It can also detect scenes featuring mascot characters in sporting event footage and select those scenes as highlights. Furthermore, it can detect scenes of audience members with pets in concert footage and select those scenes as highlights. This allows for the creation of more unique and engaging videos by detecting the movements of specific animals.
[0059] The video generation system can also be equipped with the ability to detect specific buildings and landmarks during the collection of video footage and edit the footage based on those findings. For example, it can detect churches and wedding venues in wedding footage and select those scenes as highlights. It can also detect stadiums and arenas in sporting event footage and select those scenes as highlights. Furthermore, it can detect famous halls and arenas in concert footage and select those scenes as highlights. In this way, by detecting specific buildings and landmarks, it is possible to generate more impressive and memorable videos.
[0060] The video generation system can also be equipped with the ability to detect specific seasons and times of day when collecting video footage and edit the footage based on those characteristics. For example, it can detect cherry blossoms in spring or autumn foliage in wedding footage and select those scenes as highlights. It can also detect summer blue skies or winter snowscapes in sporting event footage and select those scenes as highlights. Furthermore, it can detect nighttime lighting or sunset scenes in concert footage and select those scenes as highlights. By detecting specific seasons and times of day, it is possible to generate videos that more strongly convey a sense of the seasons and the passage of time.
[0061] The video generation system can also be equipped with the ability to detect specific clothing or fashion items during the collection of video footage and edit the video based on that detection. For example, it can detect the bride and groom's attire in wedding footage and select those scenes as highlights. It can also detect athletes' uniforms in sporting event footage and select those scenes as highlights. Furthermore, it can detect artists' costumes in concert footage and select those scenes as highlights. This allows for the creation of more visually appealing videos by detecting specific clothing or fashion items.
[0062] The following briefly describes the processing flow for example form 1.
[0063] Step 1: The editorial team sets up multiple cameras with microphones, and AI instantly identifies highlights and performs real-time editing. For example, multiple cameras with microphones are used to collect video footage, and the AI instantly identifies highlights. It is also possible to edit the video footage in real time using AI. The AI selects important scenes from the video footage and performs real-time editing. Step 2: The generation unit generates a complete work from the video footage edited by the editorial unit. For example, the AI generates a complete work based on the video footage edited by the editorial unit. It is also possible to combine video footage using the AI to generate a complete work. The AI selects the optimal scenes from the video footage and generates a complete work. Step 3: The authentication unit reads the photo data, performs facial recognition, and generates a video specifying the person to be included. For example, the AI analyzes the photo data and performs facial recognition. It can also use AI to recognize a specific person from the photo data and include them in the video. The AI recognizes a specific person from the photo data and includes that person in the video. Step 4: The reading unit reads video data shot on a smartphone or other device and generates an immersive video. For example, the AI analyzes video data shot on a smartphone and generates an immersive video. It is also possible to analyze video data using AI and generate an immersive video. The AI selects important scenes from the video data and generates an immersive video.
[0064] (Example of form 2) The video generation system according to an embodiment of the present invention is a system in which multiple microphone-equipped cameras are arranged, and AI instantly identifies highlights and performs real-time editing to complete a single video. This system can generate videos freely from any video material. Furthermore, by reading photo data and performing facial recognition, it is also possible to specify the people to be included in the video and generate it. In addition, video data shot with smartphones can also be read to generate more immersive videos. For example, multiple microphone-equipped cameras are placed in the venue of a wedding reception, and AI instantly identifies highlights and performs real-time editing to complete a single video by the end of the banquet. It can generate videos freely from any video material, and by reading group photos and performing facial recognition, it is also possible to specify the guests to be included in the video and generate it. In addition, video data shot by attendees can also be read to generate more immersive videos. Since it uses wireless LAN-compatible network cameras, it can be applied to events other than wedding receptions, provided the environment is suitable. This system was developed to address the current situation where securing photographers has become difficult amidst the increasing trend in the number of wedding receptions since the end of the COVID-19 pandemic. Customer demand is increasing, and by using AI to select and edit the best scenes, we can meet customer needs. This allows the video generation system to use multiple microphone-equipped cameras, with AI instantly identifying highlights and editing in real time to produce a complete video.
[0065] The video generation system according to this embodiment comprises an editing unit, a generation unit, an authentication unit, and a reading unit. The editing unit arranges multiple cameras with microphones, and the AI instantly identifies highlights and performs real-time editing. For example, the editing unit collects video material using multiple cameras with microphones, and the AI instantly identifies highlights. The editing unit can also edit video material in real time using the AI. For example, the editing unit has the AI select important scenes from the video material and edit them in real time. The generation unit generates a complete video from the video material edited by the editing unit. For example, the generation unit has the AI generate a complete video based on the video material edited by the editing unit. The generation unit can also combine video material using the AI to generate a complete video. For example, the generation unit has the AI select the optimal scenes from the video material and generate a complete video. The authentication unit reads photo data, performs facial recognition, and generates a video specifying the person to be included in the video. For example, the authentication unit has the AI analyze photo data and perform facial recognition. Furthermore, the authentication unit can use AI to recognize specific individuals from photo data and incorporate them into videos. For example, the authentication unit uses AI to recognize a specific person from photo data and incorporate that person into the video. The reading unit reads video data shot with a smartphone or other device and generates immersive videos. For example, the reading unit uses AI to analyze video data shot with a smartphone and generate immersive videos. The reading unit can also use AI to analyze video data and generate immersive videos. For example, the reading unit uses AI to select important scenes from video data and generate immersive videos. As a result, the video generation system according to this embodiment can use multiple microphone-equipped cameras, have AI instantly grasp the highlights, and edit in real time to generate a single work.
[0066] The editorial team deploys multiple cameras equipped with microphones, and AI instantly identifies highlights and performs real-time editing. Specifically, the editorial team strategically places multiple cameras with microphones at event venues and sports stadiums, with each camera collecting video and audio from a different perspective. These cameras are equipped with wide-angle lenses and zoom capabilities, allowing them to capture wide-area footage in high resolution. The collected video and audio data is transmitted in real time to a central server, where AI analyzes this data. The AI uses speech recognition technology to detect audio events such as cheers and applause, and identifies important scenes from the video. For example, in sports events, goal scenes and player interviews are selected as highlights. Furthermore, the AI analyzes movement and color changes in the video to select visually appealing scenes. This allows the editorial team to instantly extract highlights from a vast amount of video data and edit them in real time. The edited video is immediately distributed to viewers, providing an immersive live experience. In addition, the editorial team utilizes the AI's learning capabilities to optimize the editing algorithm based on past editing data, enabling more accurate highlight selection. This allows the editorial team to achieve efficient and high-quality video editing, enabling them to provide compelling content to viewers.
[0067] The generation unit creates a complete video from video footage edited by the editorial unit. Specifically, the generation unit uses AI to combine video footage based on highlight scenes selected by the editorial unit to create a complete video. The AI considers the flow of the video and storytelling elements, optimizing the order of scenes and transitions. For example, in a sports event highlight video, it arranges important moments of the game in chronological order and appropriately inserts audience reactions and player interviews to create a video that resonates with viewers. Furthermore, the generation unit also performs color correction and audio adjustment to improve the overall quality. The AI automatically adjusts the brightness and contrast of the video and removes audio noise to achieve clear sound quality. The generation unit can also add effects and subtitles to the video. For example, it can insert subtitles displaying player names and scores to make the information easy for viewers to understand. In this way, the generation unit can create a visually and aurally superior video based on the material provided by the editorial unit. Moreover, the generation unit utilizes the AI's learning function to optimize the generation algorithm based on past video data, enabling it to efficiently produce higher quality videos. This allows the production unit to provide viewers with consistent, high-quality video content, thereby increasing their satisfaction.
[0068] The authentication unit reads photo data, performs facial recognition, and generates videos featuring the person to be included. Specifically, the authentication unit uses AI to analyze user-provided photo data and facial recognition technology to identify a specific person. The AI extracts facial feature points and matches them with existing facial data in the database to identify the person. For example, facial photos of event participants or athletes can be registered in advance, and the system can automatically detect scenes in the video in which those individuals appear. The authentication unit can understand which scenes feature a specific person and prioritize editing and generating those scenes. This allows users to easily create video works centered around themselves or specific individuals. Furthermore, the authentication unit can use facial recognition technology to analyze the expressions and movements of people in the video and capture changes in emotion. For example, by emphasizing the expressions of joy or frustration of athletes, it can create videos that move viewers. In addition, the authentication unit can continuously improve the accuracy of facial recognition by utilizing the AI's learning function. This allows the authentication unit to accurately identify specific individuals requested by the user and include them in the video work.
[0069] The loading unit reads video data shot on smartphones and other devices and generates immersive videos. Specifically, the loading unit uses AI to analyze video data shot by the user on their smartphone and selects important scenes from the footage. The AI analyzes the movement and changes in sound in the video, detecting visually appealing scenes and audio events. For example, in travel videos, it selects scenes of beautiful scenery and enjoyable activities to create an immersive video. Furthermore, the loading unit also performs color correction and audio adjustment to improve the overall quality. The AI automatically adjusts the brightness and contrast of the video and removes audio noise to achieve clear sound quality. The loading unit can also add effects and subtitles to the video. For example, it can insert subtitles displaying place names and dates into travel videos, making it easier for viewers to understand the information. In this way, the loading unit can generate visually and aurally superior, immersive videos based on video data shot by the user. Moreover, the loading unit utilizes the AI's learning function to optimize the analysis algorithm based on past video data, enabling it to efficiently generate higher quality videos. This allows the loading unit to provide users with consistent, high-quality video content, thereby increasing their satisfaction.
[0070] The video generation system can utilize Wi-Fi-enabled network cameras. Using Wi-Fi-enabled network cameras increases flexibility in installation locations, enabling use in a variety of environments. For example, Wi-Fi-enabled network cameras can be installed at wedding reception venues. They can also be installed at venues for corporate events, sporting events, concerts, and other similar events. Wi-Fi-enabled network cameras have specific specifications, such as being from a particular manufacturer, model, or communication speed. This flexibility in installation locations and use in diverse environments is a key advantage of using Wi-Fi-enabled network cameras.
[0071] The video generation system can be applied to events other than wedding receptions. Because it can be applied to events other than wedding receptions, it can be used for a wide range of purposes. For example, the video generation system can be applied to various events such as corporate events, sporting events, and concerts. In corporate events, multiple cameras with microphones can be placed, and the AI can instantly identify highlights and edit them in real time to create a single video. In sporting events, multiple cameras with microphones can be placed, and the AI can instantly identify highlights and edit them in real time to create a single video by the end of the match. In concerts, multiple cameras with microphones can be placed, and the AI can instantly identify highlights and edit them in real time to create a single video by the end of the concert. As such, the video generation system can be applied to events other than wedding receptions and can be used for a wide range of purposes.
[0072] The editorial team can collect video footage using multiple cameras equipped with microphones and instantly identify highlights. For example, the editorial team can collect video footage using multiple cameras equipped with microphones. For instance, the editorial team can place multiple cameras equipped with microphones at a wedding reception venue to collect video footage. They can also place multiple cameras equipped with microphones at venues such as corporate events, sporting events, and concerts to collect video footage. The editorial team uses AI to instantly identify highlights from the video footage. For example, the editorial team can use AI to select important scenes from the video footage and instantly identify the highlights. This allows for the collection of video footage using multiple cameras equipped with microphones and the instant identification of highlights.
[0073] The generation unit can generate a complete work from video footage edited by the editorial unit. For example, the generation unit uses AI to generate a complete work based on video footage edited by the editorial unit. For instance, the generation unit uses AI to generate a complete work based on video footage of a wedding reception. The generation unit can also use AI to generate a complete work based on video footage of corporate events, sporting events, concerts, etc. The generation unit uses AI to combine video footage to create a complete work. For example, the generation unit uses AI to select the optimal scenes from the video footage and generate a complete work. In this way, a complete work can be generated from edited video footage.
[0074] The authentication unit can read photo data, perform facial recognition, and generate a video featuring the person to be included. For example, the authentication unit can use AI to analyze photo data and perform facial recognition. For instance, it can read a group photo from a wedding reception, use AI to perform facial recognition, and generate a video featuring the guests to be included. The authentication unit can also read photo data from corporate events, sporting events, concerts, etc., use AI to perform facial recognition, and generate a video featuring the person to be included. The authentication unit uses AI to recognize specific individuals from photo data and incorporate them into the video. For example, the authentication unit uses AI to recognize a specific person from photo data and incorporate that person into the video. This allows for facial recognition using photo data and the inclusion of specific individuals in the video.
[0075] The reading unit can read video data shot on smartphones and other devices and generate immersive videos. For example, the reading unit uses AI to analyze video data shot on a smartphone and generate immersive videos. For instance, the reading unit can read video data shot by wedding reception attendees, analyze it with AI, and generate immersive videos. The reading unit can also read video data shot at corporate events, sporting events, concerts, etc., and analyze it with AI to generate immersive videos. The reading unit uses AI to analyze video data and generate immersive videos. For example, the reading unit's AI selects important scenes from the video data and generates immersive videos. This makes it possible to generate immersive videos using video data shot on smartphones and other devices.
[0076] The editorial team can estimate the user's emotions and adjust the highlight selection criteria based on those emotions. For example, if the user is emotional, the editorial team will prioritize selecting emotionally moving scenes as highlights. For instance, they might prioritize emotionally moving scenes in wedding footage. Similarly, if the user is enjoying themselves, the editorial team can select scenes with many smiles and laughter as highlights. For example, they might select scenes with many smiles and laughter in sporting event footage. Furthermore, if the user is feeling tense, the editorial team can select relaxing scenes as highlights. For example, they might select relaxing scenes in concert footage. By adjusting the highlight selection criteria based on the user's emotions, the editorial team can select highlights that are more relevant to the user. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the editorial department may be performed using AI, for example, or without AI. For example, the editorial department can input user facial expression data into a generating AI and have the generating AI perform emotion estimation.
[0077] The editorial team can analyze audio data when collecting video footage and prioritize selecting scenes containing specific keywords as highlights. For example, the editorial team can select scenes containing keywords such as "congratulations" or "I love you" in wedding speeches as highlights. The editorial team can also select scenes containing keywords such as "goal" or "victory" in sporting events as highlights. The editorial team can also select scenes containing keywords such as "encore" or "thank you" in concerts as highlights. By analyzing audio data, the editorial team can prioritize selecting scenes containing specific keywords as highlights. Some or all of the above processing by the editorial team may be performed using AI, for example, or without AI. For example, the editorial department can input audio data into a generative AI and have the AI detect specific keywords.
[0078] The editorial team can detect specific actions or gestures when collecting video footage and select highlights based on them. For example, the editorial team could select a scene of the bride and groom kissing at a wedding as a highlight. The editorial team could also select a scene of an athlete making a victory pose at a sporting event as a highlight. The editorial team could also select a scene of an artist waving to the audience at a concert as a highlight. In this way, by detecting specific actions or gestures, highlights can be selected based on them. Some or all of the above processing in the editorial team may be performed using AI, for example, or not using AI. For example, the editorial team could input video data into a generating AI and have the generating AI perform the detection of specific actions or gestures.
[0079] The editorial team can estimate the user's emotions and adjust the display order of highlights based on the estimated emotions. For example, if the user is moved, the editorial team can display emotional scenes first. For example, in wedding footage, the editorial team can display emotional scenes first. Also, if the user is having fun, the editorial team can display fun scenes first. For example, in sporting event footage, the editorial team can display relaxing scenes first. Also, if the user is feeling tense, the editorial team can display relaxing scenes first. For example, in concert footage, the editorial team can display relaxing scenes first. By adjusting the display order of highlights based on the user's emotions, highlights can be displayed in an order that is more appropriate for the user. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the editorial team may be performed using AI, for example, or without AI. For example, the editorial department can input user facial expression data into a generative AI and have the AI perform emotion estimation.
[0080] The editorial team can analyze ambient and background sounds when collecting video footage and select scenes containing specific sounds as highlights. For example, the editorial team can select scenes with a lot of applause and cheering at a wedding as highlights. The editorial team can also select scenes with loud crowd cheering at a sporting event as highlights. The editorial team can also select scenes with audience singing at a concert as highlights. By analyzing ambient and background sounds, the editorial team can select scenes containing specific sounds as highlights. Some or all of the above processing by the editorial team may be performed using AI, for example, or without AI. For example, the editorial team can input audio data into a generating AI and have the generating AI perform the detection of specific sounds.
[0081] The editorial department can analyze the color tone and brightness of video footage during collection and select scenes with specific color tones and brightness levels as highlights. For example, the editorial department can select bright and vibrant scenes as highlights at a wedding. For example, the editorial department can select scenes where the vibrant uniforms stand out at a sporting event as highlights. For example, the editorial department can select scenes with beautiful lighting at a concert as highlights. In this way, by analyzing the color tone and brightness of the video, scenes with specific color tones and brightness levels can be selected as highlights. Some or all of the above processing by the editorial department may be performed using AI, for example, or without AI. For example, the editorial department can input video data into a generating AI and have the generating AI perform the detection of specific color tones and brightness levels.
[0082] The generation unit can estimate the user's emotions and adjust the style of the generated work based on the estimated user emotions. For example, if the user is moved, the generation unit can generate a work using emotional music and effects. For example, the generation unit can generate a work using emotional music and effects with wedding video footage. Also, if the user is having fun, the generation unit can generate a work using bright and cheerful music and effects. For example, the generation unit can generate a work using bright and cheerful music and effects with sporting event video footage. Also, if the user is feeling tense, the generation unit can generate a work using relaxing music and effects. For example, the generation unit can generate a work using relaxing music and effects with concert video footage. In this way, by adjusting the style of the generated work based on the user's emotions, it is possible to generate works that are more suitable for the user. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input user facial expression data into the generation AI and have the generation AI perform emotion estimation.
[0083] The generation unit can construct a natural storyline by considering the temporal flow of the video material during generation. For example, the generation unit can edit wedding video material in chronological order to construct a natural storyline. The generation unit can also edit sports event video material in accordance with the progress of the game to construct a natural storyline. The generation unit can also edit concert video material according to the setlist to construct a natural storyline. By considering the temporal flow of the video material, a natural storyline can be constructed. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input video data into a generation AI and have the generation AI perform editing that considers the temporal flow.
[0084] The generation unit can optimize the synchronization of audio and video in video footage during generation. For example, the generation unit can synchronize the audio and video of speeches in wedding video footage. The generation unit can also synchronize the audio and video of commentary in sporting event video footage. The generation unit can also synchronize the audio and video of performances in concert video footage. By optimizing the synchronization of audio and video in video footage, a more cohesive work can be produced. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input audio and video data into a generation AI and have the generation AI perform the synchronization optimization.
[0085] The generation unit can estimate the user's emotions and adjust the length of the generated content based on the estimated emotions. For example, if the user is in a hurry, the generation unit can generate a short, concise content. For example, it can generate a short, concise content using wedding video footage. If the user is relaxed, the generation unit can also generate a longer content with detailed explanations. For example, it can generate a longer content with detailed explanations using sporting event video footage. If the user is excited, the generation unit can also generate content with visually stimulating effects. For example, it can generate content with visually stimulating effects using concert video footage. By adjusting the length of the generated content based on the user's emotions, it is possible to generate content that is more suitable for the user. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generation AI. The generation AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input user facial expression data into the generation AI and have the generation AI perform emotion estimation.
[0086] The generation unit can create visually appealing works by combining different camera angles from video footage during the generation process. For example, the generation unit can combine different angles of the bride and groom from wedding video footage to create a work. The generation unit can also combine different angles of athletes from sporting event video footage to create a work. The generation unit can also combine different angles of artists from concert video footage to create a work. By combining different camera angles, it is possible to create visually appealing works. Some or all of the above-described processes in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input video data into a generation AI and have the generation AI execute combinations of different camera angles.
[0087] The generation unit can generate smooth video by adjusting the different frame rates of video material during generation. For example, the generation unit can generate smooth video by adjusting the different frame rates of wedding video material. The generation unit can also generate smooth video by adjusting the different frame rates of sporting event video material. For example, the generation unit can generate smooth video by adjusting the different frame rates of concert video material. In this way, smooth video can be generated by adjusting the different frame rates. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input video data into a generation AI and have the generation AI perform the frame rate adjustment.
[0088] The authentication unit can estimate the user's emotions and adjust the accuracy of authentication based on the estimated emotions. For example, if the user is nervous, the authentication unit can increase the accuracy of authentication to ensure successful authentication. For example, if the user is nervous in wedding video footage, the authentication unit can increase the accuracy of authentication to ensure successful authentication. The authentication unit can also adjust the accuracy of authentication appropriately to ensure smooth authentication if the user is relaxed. For example, if the user is relaxed in sporting event video footage, the authentication unit can adjust the accuracy of authentication appropriately to ensure smooth authentication. The authentication unit can also adjust the accuracy of authentication to ensure quick authentication if the user is in a hurry. For example, if the user is in a hurry in concert video footage, the authentication unit can adjust the accuracy of authentication to ensure quick authentication. In this way, by adjusting the accuracy of authentication based on the user's emotions, more appropriate authentication can be performed. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the authentication unit may be performed using AI, for example, or without AI. For example, the authentication unit can input user facial expression data into a generating AI and have the generating AI perform emotion estimation.
[0089] The authentication unit can apply the optimal facial recognition algorithm during authentication, taking into account the resolution of the photo data. For example, the authentication unit can perform detailed facial recognition using high-resolution photo data. For example, the authentication unit can perform detailed facial recognition using high-resolution photo data of a wedding. The authentication unit can also perform simple facial recognition using low-resolution photo data. For example, the authentication unit can perform simple facial recognition using low-resolution photo data of a sporting event. The authentication unit can also perform balanced facial recognition using medium-resolution photo data. For example, the authentication unit can perform balanced facial recognition using medium-resolution photo data of a concert. This improves authentication accuracy by applying the optimal facial recognition algorithm according to the resolution of the photo data. Some or all of the above processing in the authentication unit may be performed using AI, for example, or without AI. For example, the authentication unit can input photo data into a generating AI and have the generating AI perform facial recognition according to the resolution.
[0090] The authentication unit can perform more accurate facial recognition by combining multiple photographic data during authentication. For example, the authentication unit can perform facial recognition by combining photographic data taken from multiple angles. For example, the authentication unit can perform facial recognition by combining photographic data taken from multiple angles at a wedding. The authentication unit can also perform facial recognition by combining photographic data taken at multiple different times. For example, the authentication unit can perform facial recognition by combining photographic data taken at multiple different times at a sporting event. The authentication unit can also perform facial recognition by combining photographic data taken at multiple different locations. For example, the authentication unit can perform facial recognition by combining photographic data taken at multiple different locations at a concert. By combining multiple photographic data, more accurate facial recognition can be achieved. Some or all of the above processing in the authentication unit may be performed using AI, for example, or without AI. For example, the authentication unit can input multiple photographic data into a generating AI and have the generating AI perform facial recognition.
[0091] The authentication unit can estimate the user's emotions and adjust the display method of the authentication result based on the estimated emotions. For example, if the user is nervous, the authentication unit can provide a simple and easy-to-read display method. For example, if the user is nervous watching wedding video footage, the authentication unit can provide a simple and easy-to-read display method. The authentication unit can also provide a display method that includes detailed information if the user is relaxed. For example, if the user is relaxed watching sporting event video footage, the authentication unit can provide a display method that includes detailed information. The authentication unit can also provide a concise display method if the user is in a hurry. For example, if the user is in a hurry watching concert video footage, the authentication unit can provide a concise display method. In this way, by adjusting the display method of the authentication result based on the user's emotions, a more appropriate display method can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the authentication unit may be performed using AI, for example, or without AI. For example, the authentication unit can input user facial expression data into a generating AI and have the generating AI perform emotion estimation.
[0092] The authentication unit can prioritize the use of the most recent photograph during authentication, taking into account the date and time the photograph was taken. For example, the authentication unit can perform facial recognition using the most recent photograph. For example, the authentication unit can perform facial recognition using the most recent photograph from a wedding. The authentication unit can also perform facial recognition using past photographs. For example, the authentication unit can perform facial recognition using past photographs from a sporting event. The authentication unit can also perform facial recognition using photographs taken within a specific period. For example, the authentication unit can perform facial recognition using photographs taken within a specific period from a concert. By considering the date and time the photograph was taken, the authentication unit can prioritize the use of the most recent photograph, thereby improving authentication accuracy. Some or all of the above processing in the authentication unit may be performed using AI, for example, or without AI. For example, the authentication unit can input photographic data into a generating AI and have the generating AI perform facial recognition based on the date and time the photograph was taken.
[0093] The authentication unit analyzes the background information of the photo data during authentication and can prioritize the use of photos that contain a specific background. For example, the authentication unit can perform facial recognition using photo data taken at a specific location. For example, it can perform facial recognition using photo data taken at a specific location at a wedding. The authentication unit can also perform facial recognition using photo data taken at a specific event. For example, it can perform facial recognition using photo data taken at a specific sporting event. The authentication unit can also perform facial recognition using photo data taken at a specific time of day. For example, it can perform facial recognition using photo data taken at a specific time of day at a concert. By analyzing the background information of the photo data, the authentication unit can prioritize the use of photos that contain a specific background, thereby improving authentication accuracy. Some or all of the above processing in the authentication unit may be performed using AI, for example, or without AI. For example, the authentication unit can input photo data into a generating AI and have the generating AI perform facial recognition based on the background information.
[0094] The loading unit can estimate the user's emotions and determine the priority of video data to load based on the estimated emotions. For example, if the user is emotional, the loading unit will prioritize loading video data containing emotional scenes. For example, it will prioritize loading video data containing emotional scenes from wedding footage. Also, if the user is having fun, the loading unit can prioritize loading video data containing fun scenes from sporting event footage. Furthermore, if the user is feeling tense, the loading unit can prioritize loading video data containing relaxing scenes. For example, it will prioritize loading video data containing relaxing scenes from concert footage. In this way, by prioritizing video data based on the user's emotions, more appropriate video data can be loaded preferentially. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the reading unit may be performed using AI, for example, or without AI. For example, the reading unit can input user facial expression data into a generating AI and have the generating AI perform emotion estimation.
[0095] The loading unit can select the optimal loading method considering the resolution of the video data during loading. For example, the loading unit can prioritize loading high-resolution video data to provide detailed footage. For example, the loading unit can prioritize loading high-resolution video data of a wedding to provide detailed footage. The loading unit can also prioritize loading low-resolution video data to enable faster playback. For example, the loading unit can prioritize loading low-resolution video data of a sporting event to enable faster playback. The loading unit can also prioritize loading medium-resolution video data to provide balanced footage. For example, the loading unit can prioritize loading medium-resolution video data of a concert to provide balanced footage. In this way, by selecting the optimal loading method according to the resolution of the video data, higher quality video can be provided. Some or all of the above processing in the loading unit may be performed using AI, for example, or without AI. For example, the loading unit can input video data into a generating AI and have the generating AI perform loading according to the resolution.
[0096] The loading unit can adjust the frame rate of video data during loading to achieve smooth playback. For example, the loading unit can prioritize loading high-frame-rate video data to achieve smooth playback. For example, the loading unit can prioritize loading high-frame-rate video data of a wedding to achieve smooth playback. The loading unit can also prioritize loading low-frame-rate video data to achieve fast playback. For example, the loading unit can prioritize loading low-frame-rate video data of a sporting event to achieve fast playback. The loading unit can also prioritize loading medium-frame-rate video data to achieve balanced playback. For example, the loading unit can prioritize loading medium-frame-rate video data of a concert to achieve balanced playback. In this way, smooth playback can be achieved by adjusting the frame rate of the video data. Some or all of the above processing in the loading unit may be performed using AI, for example, or without AI. For example, the loading unit can input video data into a generating AI and have the generating AI perform the frame rate adjustment.
[0097] The loading unit can estimate the user's emotions and adjust how the loaded video data is displayed based on the estimated emotions. For example, if the user is moved, the loading unit can highlight emotional scenes. For example, it can highlight emotional scenes in wedding video footage. The loading unit can also highlight enjoyable scenes if the user is having fun. For example, it can highlight enjoyable scenes in sporting event video footage. The loading unit can also highlight relaxing scenes if the user is feeling tense. For example, it can highlight relaxing scenes in concert video footage. By adjusting how the video data is displayed based on the user's emotions, a more appropriate display method can be provided. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the loading unit may be performed using AI, for example, or without AI. For example, the reading unit can input the user's facial expression data into a generating AI, allowing the generating AI to perform emotion estimation.
[0098] The loading unit can prioritize loading the latest video data by considering the shooting date and time of the video data during loading. For example, the loading unit can prioritize loading the latest video data and provide the most recent footage. For example, the loading unit can prioritize loading the latest video data of a wedding and provide the most recent footage. The loading unit can also prioritize loading past video data and provide past footage. For example, the loading unit can prioritize loading past video data of a sporting event and provide past footage. The loading unit can also prioritize loading video data shot within a specific period and provide footage from that period. For example, the loading unit can prioritize loading video data shot within a specific period of a concert and provide footage from that period. This allows the loading unit to prioritize loading the latest video data by considering the shooting date and time of the video data. Some or all of the above processing in the loading unit may be performed using AI, for example, or without AI. For example, the loading unit can input video data into a generating AI and have the generating AI perform loading based on the shooting date and time.
[0099] The loading unit can analyze the audio information of video data during loading and prioritize loading videos that contain specific audio. For example, the loading unit can prioritize loading video data that contains audio with specific keywords. For example, it can prioritize loading video data from weddings that contains audio with keywords such as "congratulations" or "I love you." The loading unit can also prioritize loading video data that contains specific music. For example, it can prioritize loading video data from sporting events that contains specific music. The loading unit can also prioritize loading video data that contains specific ambient sounds. For example, it can prioritize loading video data from concerts that contain specific ambient sounds. In this way, by analyzing the audio information of video data, it is possible to prioritize loading videos that contain specific audio. Some or all of the above processing in the loading unit may be performed using AI, for example, or without AI. For example, the loading unit can input audio data into a generating AI and have the generating AI perform the detection of specific audio.
[0100] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0101] The video generation system can also be equipped with a function to estimate the user's emotions and automatically adjust the video's color tone based on those emotions. For example, if the user is moved, the video's color tone can be adjusted to warmer tones to emphasize the emotional atmosphere. If the user is having fun, the video's color tone can be adjusted to be brighter and more vivid to emphasize the enjoyable atmosphere. Furthermore, if the user is relaxed, the video's color tone can be adjusted to softer, calmer tones to emphasize the relaxed atmosphere. In this way, by automatically adjusting the video's color tone based on the user's emotions, it is possible to generate more emotionally appealing videos.
[0102] The video generation system can also be equipped with the ability to detect specific weather conditions when collecting video footage and automatically add video effects based on those conditions. For example, during rainy weather, a rain effect can be added to the video to enhance realism. During sunny weather, a sunlight effect can be added to create a bright and refreshing atmosphere. Furthermore, on snowy days, a snow effect can be added to emphasize the winter atmosphere. In this way, by automatically adding video effects based on weather conditions, it is possible to generate more realistic and immersive videos.
[0103] The video generation system can also be equipped with the ability to estimate the user's emotions and automatically select the audio track for the video based on those emotions. For example, if the user is moved, emotional music can be selected and added to the video. If the user is having fun, bright and cheerful music can be selected and added to the video. Furthermore, if the user is relaxed, calming music can be selected and added to the video. In this way, by automatically selecting the audio track based on the user's emotions, it is possible to generate more emotionally appealing videos.
[0104] The video generation system can also be equipped with the ability to detect the movements of specific animals during the collection of video footage and edit the footage based on that detection. For example, it can detect scenes featuring pets in wedding footage and select those scenes as highlights. It can also detect scenes featuring mascot characters in sporting event footage and select those scenes as highlights. Furthermore, it can detect scenes of audience members with pets in concert footage and select those scenes as highlights. This allows for the creation of more unique and engaging videos by detecting the movements of specific animals.
[0105] The video generation system can also be equipped with the ability to estimate the user's emotions and automatically select transition effects for the video based on those emotions. For example, if the user is moved, it can select soft transition effects such as fade-in or fade-out. If the user is enjoying themselves, it can select dynamic transition effects such as slides or zooms. Furthermore, if the user is relaxed, it can select smooth transition effects such as dissolves or crossfades. By automatically selecting transition effects based on the user's emotions, it is possible to generate videos that appeal more to emotions.
[0106] The video generation system can also be equipped with the ability to detect specific buildings and landmarks during the collection of video footage and edit the footage based on those findings. For example, it can detect churches and wedding venues in wedding footage and select those scenes as highlights. It can also detect stadiums and arenas in sporting event footage and select those scenes as highlights. Furthermore, it can detect famous halls and arenas in concert footage and select those scenes as highlights. In this way, by detecting specific buildings and landmarks, it is possible to generate more impressive and memorable videos.
[0107] The video generation system can also be equipped with the ability to estimate the user's emotions and automatically adjust the playback speed of the video based on those emotions. For example, if the user is moved, slow motion can be used to emphasize the emotional scenes. If the user is having fun, enjoyable scenes can be played at normal speed. Furthermore, if the user is feeling tense, fast-forward can be used to play scenes that alleviate the tension. In this way, by automatically adjusting the playback speed based on the user's emotions, it is possible to generate more emotionally resonant videos.
[0108] The video generation system can also be equipped with the ability to detect specific seasons and times of day when collecting video footage and edit the footage based on those characteristics. For example, it can detect cherry blossoms in spring or autumn foliage in wedding footage and select those scenes as highlights. It can also detect summer blue skies or winter snowscapes in sporting event footage and select those scenes as highlights. Furthermore, it can detect nighttime lighting or sunset scenes in concert footage and select those scenes as highlights. By detecting specific seasons and times of day, it is possible to generate videos that more strongly convey a sense of the seasons and the passage of time.
[0109] The video generation system can also be equipped with the ability to estimate the user's emotions and automatically generate subtitles based on those emotions. For example, if the user is moved, an emotional message can be displayed as subtitles. If the user is enjoying themselves, fun comments or jokes can be displayed as subtitles. Furthermore, if the user is relaxed, a relaxing message can be displayed as subtitles. By automatically generating subtitles based on the user's emotions, it becomes possible to create more emotionally resonant videos.
[0110] The video generation system can also be equipped with the ability to detect specific clothing or fashion items during the collection of video footage and edit the video based on that detection. For example, it can detect the bride and groom's attire in wedding footage and select those scenes as highlights. It can also detect athletes' uniforms in sporting event footage and select those scenes as highlights. Furthermore, it can detect artists' costumes in concert footage and select those scenes as highlights. This allows for the creation of more visually appealing videos by detecting specific clothing or fashion items.
[0111] The following briefly describes the processing flow for example form 2.
[0112] Step 1: The editorial team sets up multiple cameras with microphones, and AI instantly identifies highlights and performs real-time editing. For example, multiple cameras with microphones are used to collect video footage, and the AI instantly identifies highlights. It is also possible to edit the video footage in real time using AI. The AI selects important scenes from the video footage and performs real-time editing. Step 2: The generation unit generates a complete work from the video footage edited by the editorial unit. For example, the AI generates a complete work based on the video footage edited by the editorial unit. It is also possible to combine video footage using the AI to generate a complete work. The AI selects the optimal scenes from the video footage and generates a complete work. Step 3: The authentication unit reads the photo data, performs facial recognition, and generates a video specifying the person to be included. For example, the AI analyzes the photo data and performs facial recognition. It can also use AI to recognize a specific person from the photo data and include them in the video. The AI recognizes a specific person from the photo data and includes that person in the video. Step 4: The reading unit reads video data shot on a smartphone or other device and generates an immersive video. For example, the AI analyzes video data shot on a smartphone and generates an immersive video. It is also possible to analyze video data using AI and generate an immersive video. The AI selects important scenes from the video data and generates an immersive video.
[0113] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0114] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0115] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0116] Each of the multiple elements described above, including the editing unit, generation unit, authentication unit, and reading unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the editing unit collects video material using the camera 42 and microphone 38B of the smart device 14, and the control unit 46A uses AI to instantly identify highlights and perform editing in real time. The generation unit generates a complete video from the edited video material using the specific processing unit 290 of the data processing unit 12. The authentication unit analyzes photo data using the specific processing unit 290 of the data processing unit 12, performs facial recognition, and generates a video with a specified person to be included. The reading unit analyzes video data shot with a smartphone using the control unit 46A of the smart device 14 and generates an immersive video. The correspondence between each unit and the device or control unit is not limited to the examples described above and can be modified in various ways.
[0117] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0118] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0119] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0120] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0121] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0122] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0123] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0124] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0125] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0126] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0127] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0128] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0129] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0130] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0131] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0132] Each of the multiple elements described above, including the editing unit, generation unit, authentication unit, and reading unit, is implemented, for example, in at least one of the smart glasses 214 and the data processing unit 12. For example, the editing unit collects video material using the camera 42 and microphone 238 of the smart glasses 214, and the control unit 46A uses AI to instantly grasp highlights and perform editing in real time. The generation unit generates a complete video from the edited video material using, for example, the identification processing unit 290 of the data processing unit 12. The authentication unit analyzes photo data using, for example, the identification processing unit 290 of the data processing unit 12, performs facial recognition, and generates a video with a specified person to be included. The reading unit analyzes video data shot with a smartphone using, for example, the control unit 46A of the smart glasses 214 and generates an immersive video. The correspondence between each unit and the device or control unit is not limited to the examples described above and can be modified in various ways.
[0133] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0134] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0135] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0136] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0137] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0138] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0139] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0140] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0141] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0142] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0143] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0144] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0145] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0146] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0147] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0148] Each of the multiple elements described above, including the editing unit, generation unit, authentication unit, and reading unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the editing unit collects video material using the camera 42 and microphone 238 of the headset terminal 314, and the control unit 46A uses AI to instantly grasp highlights and perform editing in real time. The generation unit generates a complete video from the edited video material using the specific processing unit 290 of the data processing unit 12. The authentication unit analyzes photo data using the specific processing unit 290 of the data processing unit 12, performs facial recognition, and generates a video specifying the person to be included. The reading unit analyzes video data shot with a smartphone using the control unit 46A of the headset terminal 314 and generates an immersive video. The correspondence between each unit and the device or control unit is not limited to the examples described above and can be modified in various ways.
[0149] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0150] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0151] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0152] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0153] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0154] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0155] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0156] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0157] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0158] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0159] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0160] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0161] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0162] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0163] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0164] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0165] Each of the multiple elements described above, including the editing unit, generation unit, authentication unit, and reading unit, is implemented by, for example, at least one of the robot 414 and the data processing unit 12. For example, the editing unit collects video material using the camera 42 and microphone 238 of the robot 414, and the control unit 46A uses AI to instantly grasp highlights and perform editing in real time. The generation unit generates a single work from the edited video material by, for example, the specific processing unit 290 of the data processing unit 12. The authentication unit analyzes photo data by, for example, the specific processing unit 290 of the data processing unit 12, performs facial recognition, and generates a video specifying the person to be included. The reading unit analyzes video data shot with a smartphone by, for example, the control unit 46A of the robot 414 and generates an immersive video. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.
[0166] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0167] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0168] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0169] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0170] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0171] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0172] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0173] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0174] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0175] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0176] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0177] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0178] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0179] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0180] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0181] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0182] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0183] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0184] (Note 1) By placing multiple cameras with microphones, The editorial team uses AI to instantly grasp the highlights and perform real-time editing, A generation unit that generates a single work from video materials edited by the aforementioned editorial department, The authentication unit reads photo data, performs facial recognition, and generates a video by specifying the person to be included in the video. It includes a reading unit that reads video data shot on a smartphone or other device and generates immersive videos. A system characterized by the following features. (Note 2) Use a Wi-Fi enabled network camera. The system described in Appendix 1, characterized by the features described herein. (Note 3) It can also be applied to events other than wedding receptions. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned editorial department, Collect video footage using multiple cameras equipped with microphones to instantly identify highlights. The system described in Appendix 1, characterized by the features described herein. (Note 5) The generating unit is The aforementioned editorial department generates a single work from the edited video footage. The system described in Appendix 1, characterized by the features described herein. (Note 6) The authentication unit, The system reads photo data, performs facial recognition, and generates a video by specifying the person to be included in the video. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned reading unit, It reads video data shot on smartphones and other devices and generates immersive videos. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned editorial department, It estimates the user's emotions and adjusts the highlight selection criteria based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned editorial department, During the collection of video footage, audio data is analyzed, and scenes containing specific keywords are prioritized and selected as highlights. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned editorial department, During the collection of video footage, specific actions and gestures are detected, and highlights are selected based on these. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned editorial department, It estimates the user's emotions and adjusts the display order of highlights based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned editorial department, During the collection of video footage, ambient and background sounds are analyzed, and scenes containing specific sounds are selected as highlights. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned editorial department, When collecting video footage, the color tone and brightness of the video are analyzed, and scenes with specific color tones and brightness levels are selected as highlights. The system described in Appendix 1, characterized by the features described herein. (Note 14) The generating unit is It estimates the user's emotions and adjusts the style of the generated work based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 15) The generating unit is During generation, the temporal flow of the video footage is taken into consideration to create a natural storyline. The system described in Appendix 1, characterized by the features described herein. (Note 16) The generating unit is During generation, the synchronization of audio and video from the video material is optimized. The system described in Appendix 1, characterized by the features described herein. (Note 17) The generating unit is It estimates the user's emotions and adjusts the length of the generated work based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 18) The generating unit is During the generation process, different camera angles from the video footage are combined to create visually appealing works. The system described in Appendix 1, characterized by the features described herein. (Note 19) The generating unit is During generation, the different frame rates of the video material are adjusted to produce smooth video. The system described in Appendix 1, characterized by the features described herein. (Note 20) The authentication unit, It estimates the user's emotions and adjusts the accuracy of authentication based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 21) The authentication unit, During authentication, the optimal facial recognition algorithm is applied, taking into account the resolution of the photo data. The system described in Appendix 1, characterized by the features described herein. (Note 22) The authentication unit, During authentication, multiple photo data are combined to perform more accurate facial recognition. The system described in Appendix 1, characterized by the features described herein. (Note 23) The authentication unit, The system estimates the user's emotions and adjusts how authentication results are displayed based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 24) The authentication unit, During authentication, the system prioritizes using the most recent photo, taking into account the date and time the photo was taken. The system described in Appendix 1, characterized by the features described herein. (Note 25) The authentication unit, During authentication, the background information of the photo data is analyzed, and photos containing specific backgrounds are given priority for use. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned reading unit, It estimates the user's emotions and determines the priority of video data to load based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned reading unit, During loading, the system selects the optimal loading method, taking into account the resolution of the video data. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned reading unit, During loading, the frame rate of the video data is adjusted to achieve smooth playback. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned reading unit, It estimates the user's emotions and adjusts how the video data is displayed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned reading unit, When loading, the system prioritizes loading the most recent video, taking into account the video data's shooting date and time. The system described in Appendix 1, characterized by the features described herein. (Note 31) The aforementioned reading unit, During loading, the system analyzes the audio information of the video data and prioritizes loading videos that contain specific audio. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]
[0185] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. By placing multiple cameras with microphones, The editorial team uses AI to instantly grasp the highlights and perform real-time editing, A generation unit that generates a single work from video materials edited by the aforementioned editorial department, The authentication unit reads photo data, performs facial recognition, and generates a video by specifying the person to be included in the video. It includes a reading unit that reads video data shot on a smartphone or other device and generates immersive videos. A system characterized by the following features.
2. Use a Wi-Fi enabled network camera. The system according to feature 1.
3. It can also be applied to events other than wedding receptions. The system according to feature 1.
4. The aforementioned editorial department, Collect video footage using multiple cameras equipped with microphones to instantly identify highlights. The system according to feature 1.
5. The generating unit is The aforementioned editorial department generates a single work from the edited video footage. The system according to feature 1.
6. The authentication unit, The system reads photo data, performs facial recognition, and generates a video by specifying the person to be included in the video. The system according to feature 1.
7. The aforementioned reading unit, It reads video data shot on smartphones and other devices and generates immersive videos. The system according to feature 1.
8. The aforementioned editorial department, It estimates the user's emotions and adjusts the highlight selection criteria based on those estimated emotions. The system according to feature 1.
9. The aforementioned editorial department, During the collection of video footage, audio data is analyzed, and scenes containing specific keywords are prioritized and selected as highlights. The system according to feature 1.
10. The aforementioned editorial department, During the collection of video footage, specific actions and gestures are detected, and highlights are selected based on these. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A