system

The system addresses real-time video editing and facial recognition challenges by using AI to analyze, edit, and generate videos with improved quality and consistency, emphasizing specific individuals and scenes.

JP2026018734APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024120062
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Conventional technologies face challenges in real-time editing of video material and generating images of specific individuals using facial recognition, necessitating improvements.

Method used

A system incorporating a video analysis unit, editing unit, face authentication unit, data reading unit, and network camera unit, utilizing AI technology to analyze, edit, and generate videos in real-time, including facial recognition and network camera integration.

Benefits of technology

Enables real-time editing and generation of videos that include specific individuals, with improved quality and consistency, emphasizing important scenes and emotions, and integrating footage from multiple sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026018734000001_ABST
    Figure 2026018734000001_ABST
Patent Text Reader

Abstract

An object of the system according to the embodiment is to edit a video material in real time and generate a video including a specific person.SOLUTION: A system includes a video analysis unit, an editing unit, a face authentication unit, a data reading unit, and a network camera unit. The video analysis unit analyzes a video material. The editing unit performs editing in real time on the basis of a result analyzed by the video analyzing unit. The face recognition unit reads a group photograph and performs face recognition. The data reading unit reads moving image data captured by a participant. The network camera unit collects video using a network camera compatible with a wireless LAN.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional technologies have difficulty in real-time editing of video material and generating images of specific individuals using facial recognition, and there is room for improvement.

[0005] The system according to the embodiment aims to edit video material in real time and generate a video including a specific person. [Means for solving the problem]

[0006] The system according to the embodiment includes a video analysis unit, an editing unit, a face authentication unit, a data reading unit, and a network camera unit. The video analysis unit analyzes video material. The editing unit performs real-time editing based on the results of analysis by the video analysis unit. The face authentication unit reads a group photo and performs face authentication. The data reading unit reads video data captured by attendees. The network camera unit collects video using a wireless LAN-compatible network camera. [Effects of the Invention]

[0007] The system according to the embodiment can edit video material in real time to generate a video including a specific person. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. DETAILED DESCRIPTION OF THE INVENTION

[0009] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0010] First, the terms used in the following description will be explained.

[0011] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, the processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (Tensor Processing Unit).

[0012] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0013] In the following embodiments, the coded storage is one or more nonvolatile storage devices that store various programs, various parameters, etc. Examples of nonvolatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0014] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), and Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0016] [First embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0020] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (for example, a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 (see FIG. 2) acquires the data indicating the user input.

[0021] Output device 40 includes a display 40A and a speaker 40B, and presents data to a user by outputting the data in a form of expression that the user can perceive (e.g., audio and / or text). Display 40A displays visible information such as text and images in accordance with instructions from processor 46. Speaker 40B outputs audio in accordance with instructions from processor 46. Camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0023] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0025] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0026] In the smart device 14, the specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. Note that the smart device 14 may have a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59.

[0027] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains a processing result (prediction result, etc.) using the data generation model 58 by communicating with the server device having the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device owned by a user (e.g., a mobile phone, a robot, a home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example 1) The video generation system according to the embodiment of the present invention is a system that uses AI technology to instantly analyze video material, edit it in real time, and generate a complete work. This allows the video generation system to analyze video material, edit it in real time, and generate an impressive work.

[0029] A video generation system according to an embodiment includes a video analysis unit, an editing unit, a facial recognition unit, a data reading unit, and a network camera unit. The video analysis unit analyzes video footage. For example, the generation AI analyzes the video footage and instantly identifies important scenes and highlights. The generation AI also analyzes the movements and gestures in the video footage and extracts scenes with significant movement as highlights. The generation AI also uses voice recognition to detect important statements and audio events and extract them as highlights. The editing unit performs real-time editing based on the results of the analysis by the video analysis unit. For example, the generation AI performs real-time editing based on the analyzed highlights to generate a complete work. The generation AI also automatically cuts and edits video to match the rhythm and tempo of the music. The generation AI also automatically adjusts the color tone and brightness in the video to generate a visually consistent work. The facial recognition unit reads a group photo and performs facial recognition. For example, the generation AI reads a group photo and identifies guests using facial recognition technology. The generation AI also analyzes the facial expressions and movements of guests and prioritizes editing scenes that highlight specific emotions. The generation AI also compares the guest's past video data to generate scenes that highlight the guest's growth and changes. The data import unit imports video data shot by attendees. For example, the generation AI imports video data shot by attendees and uses it as video material. When analyzing attendees' video data, the generation AI automatically stabilizes the video and removes noise to generate high-quality video. When analyzing attendees' video data, the generation AI also analyzes the audio in the video, extracting important comments and audio events and incorporating them into the editing. The network camera unit collects video using a wireless LAN-enabled network camera. For example, the generation AI uses a wireless LAN-enabled network camera to stream video in real time and instantly analyze and edit the video. The generation AI also uses a wireless LAN-enabled network camera to simultaneously collect video from multiple cameras and integrate and edit the video. The generative AI also uses wireless LAN-enabled network cameras to collect footage from different events and locations, then integrates and edits it.As a result, the video generation system according to the embodiment can analyze video footage, edit it in real time, and create a moving work. For example, the system can analyze video footage of a wedding reception, extract important scenes, and edit it in real time to create a moving work. It can also read video data taken by attendees and create a realistic video. It can also use a wireless LAN-compatible network camera to collect video in real time and instantly analyze and edit it.

[0030] The video analysis unit can use voice recognition to detect important statements and audio events and extract them as highlights. For example, when the generation AI analyzes video material, the video analysis unit uses voice recognition technology to detect important statements and audio events. For example, scenes with important audio content, such as wedding reception speeches and vows, can be automatically extracted. The generation AI also uses voice recognition technology to analyze audio within the video, extract important statements and audio events, and incorporate them into the editing. This makes it possible to highlight important scenes by detecting important statements and audio events using voice recognition and extracting them as highlights.

[0031] The video analysis unit analyzes movements and gestures within the video and can extract scenes with large movements as highlights. For example, when the generation AI analyzes video material, the video analysis unit analyzes movements within the video and extracts scenes with large movements as highlights. For example, scenes with a lot of movement, such as dance scenes or the moment of cutting a cake, are automatically extracted. The video analysis unit also allows the generation AI to analyze gestures within the video and extract scenes with large movements as highlights. This allows scenes with a lot of movement to be emphasized by analyzing movements and gestures within the video and extracting scenes with a lot of movement as highlights.

[0032] The video analysis unit can simultaneously analyze video from different camera angles and viewpoints to generate highlights from multiple viewpoints. For example, the video analysis unit uses a generation AI to simultaneously analyze video from different camera angles to generate highlights from multiple viewpoints. For example, the video analysis unit analyzes video from multiple cameras at a wedding reception to extract important scenes from multiple angles. The video analysis unit also uses a generation AI to analyze video from different viewpoints to generate highlights from multiple viewpoints. This allows the video analysis unit to simultaneously analyze video from different camera angles and viewpoints to generate highlights from multiple viewpoints, thereby providing a more realistic video.

[0033] The video analysis unit can compare it with past event data, detect similar scenes and patterns, and extract them as highlights. For example, when the generation AI analyzes video material, the video analysis unit compares it with past event data, detects similar scenes, and extracts them as highlights. For example, it compares it with footage of a past wedding reception to extract similar important scenes. The video analysis unit also allows the generation AI to compare it with past event data, detect similar patterns, and extract them as highlights. This makes it possible to highlight important scenes by comparing it with past event data, detecting similar scenes and patterns, and extracting them as highlights.

[0034] The editing department can automatically cut and edit video to match the rhythm and tempo of the music. For example, the generative AI automatically cuts and edits video to match the rhythm and tempo of the music during real-time editing. For example, video of a wedding reception can be cut to match the beat of the music to generate a rhythmic video. The editing department can also automatically cut and edit video to match the rhythm and tempo of the music using the generative AI. This allows for harmony between the visual and auditory senses by automatically cutting and editing video to match the rhythm and tempo of the music.

[0035] The editing department can automatically adjust the color tone and brightness within the video to generate a visually consistent work. For example, the generative AI can automatically adjust the color tone and brightness within the video during real-time editing to generate a visually consistent work. For example, the generative AI can unify the color tones of footage of a wedding reception to create a professional finish. The editing department can also have the generative AI automatically adjust the color tone and brightness within the video to generate a visually consistent work. This automatically adjusts the color tone and brightness within the video to generate a visually consistent work, allowing for a professional finish.

[0036] During real-time editing, the editing department can edit the video based on keywords and phrases specified by the user to emphasize a particular message. For example, during real-time editing, the editing department can edit the video based on keywords and phrases specified by the user to emphasize a particular message. For example, editing can be done based on keywords such as "love" or "gratitude." Furthermore, during real-time editing, the generative AI can edit the video based on keywords and phrases specified by the user to emphasize a particular message. In this way, by editing the video based on keywords and phrases specified by the user during real-time editing to emphasize a particular message, it is possible to visually convey a message.

[0037] The facial recognition unit can compare the video with past video data of guests and generate scenes that highlight the growth and changes of specific individuals. For example, when the generation AI recognizes a face, the facial recognition unit compares the video with past video data of guests and generates scenes that highlight the growth and changes of specific individuals. For example, the generation AI compares video footage of past weddings with the current wedding and edits it. The facial recognition unit also allows the generation AI to compare the video with past video data of guests and generate scenes that highlight the growth and changes of specific individuals. This allows the generation AI to compare the video with past video data of guests and generate scenes that highlight the growth and changes of specific individuals, creating a moving work.

[0038] The facial recognition unit can analyze the clothing and accessories of guests and generate video that matches a specific theme or style. For example, when the generation AI recognizes their faces, the facial recognition unit analyzes the clothing and accessories of guests and generates video that matches a specific theme or style. For example, it generates video that matches the dress code of a wedding reception. The facial recognition unit also allows the generation AI to analyze the clothing and accessories of guests and generate video that matches a specific theme or style. This makes it possible to generate visually consistent works by analyzing the clothing and accessories of guests and generating video that matches a specific theme or style.

[0039] The facial recognition unit can analyze the relationships between guests and prioritize editing scenes based on those relationships. For example, when the generation AI recognizes their faces, the facial recognition unit analyzes the relationships between guests and prioritizes editing scenes based on those relationships. For example, scenes with family and friends are prioritized. The facial recognition unit also allows the generation AI to analyze the relationships between guests and prioritize editing scenes based on those relationships. This makes it possible to create moving works by analyzing the relationships between guests and prioritize editing scenes based on those relationships.

[0040] The data reading unit can automatically stabilize the video and remove noise when analyzing the attendees' video data, thereby generating high-quality video. For example, the data reading unit can automatically stabilize the video and generate high-quality video when the generation AI analyzes the attendees' video data. For example, the data reading unit can perform image stabilization to generate stable video. Furthermore, the data reading unit can automatically remove noise when the generation AI analyzes the attendees' video data, thereby generating high-quality video. As a result, the data reading unit can automatically stabilize the video and remove noise when analyzing the attendees' video data, thereby generating high-quality video, thereby achieving a professional finish.

[0041] When analyzing attendees' video data, the data loading unit analyzes the audio in the video, extracts important remarks and audio events, and incorporates them into the editing. For example, when the generation AI analyzes attendees' video data, the data loading unit analyzes the audio in the video, extracts important remarks and audio events, and incorporates them into the editing. For example, speeches and moving words are extracted and edited. The data loading unit also enables the generation AI to analyze the audio in the video, extract important remarks and audio events, and incorporate them into the editing. This allows important scenes to be emphasized by analyzing the audio in the video, extracting important remarks and audio events, and incorporating them into the editing.

[0042] When analyzing the video data of attendees, the data loading unit can combine footage from different viewpoints and angles to generate highly immersive video. For example, when the generation AI analyzes the video data of attendees, the data loading unit can combine footage from different viewpoints and angles to generate highly immersive video. For example, the data loading unit can combine and edit footage shot by multiple attendees. In addition, the data loading unit can allow the generation AI to combine footage from different viewpoints and angles to generate highly immersive video. As a result, when analyzing the video data of attendees, the data loading unit can combine footage from different viewpoints and angles to generate highly immersive video, making it possible to create visually appealing works.

[0043] The data loading unit can edit the video based on a specific theme or storyline when analyzing the video data of attendees, thereby generating a consistent work. For example, when the generation AI analyzes the video data of attendees, the data loading unit can edit the video based on a specific theme or storyline to generate a consistent work. For example, editing based on a wedding reception theme. The data loading unit can also edit the video based on a specific theme or storyline when the generation AI analyzes the video data of attendees, thereby generating a consistent work. In this way, when analyzing the video data of attendees, the data loading unit can edit the video based on a specific theme or storyline to generate a consistent work, thereby generating a visually appealing work.

[0044] The network camera unit uses a wireless LAN-enabled network camera to stream video in real time, and the generation AI can instantly analyze and edit that video. The network camera unit, for example, uses a wireless LAN-enabled network camera to stream video in real time, and the generation AI can instantly analyze and edit that video. For example, video of a wedding reception can be analyzed in real time to extract important scenes. The network camera unit also uses a wireless LAN-enabled network camera to stream video in real time, and the generation AI can instantly analyze and edit that video. This makes it possible to stream video in real time using a wireless LAN-enabled network camera, and the generation AI can instantly analyze and edit that video, creating impressive works in real time.

[0045] The network camera unit uses wireless LAN-enabled network cameras to simultaneously collect footage from multiple cameras, and the generation AI can integrate and edit it. The network camera unit, for example, uses wireless LAN-enabled network cameras to simultaneously collect footage from multiple cameras, and the generation AI can integrate and edit it. For example, footage from multiple cameras of a wedding reception can be integrated and edited. The network camera unit also uses wireless LAN-enabled network cameras to simultaneously collect footage from multiple cameras, and the generation AI can integrate and edit it. This allows for the creation of a work that is full of realism.

[0046] The network camera unit uses a wireless LAN-enabled network camera to collect footage from different events and locations, and the generation AI can integrate and edit them. The network camera unit, for example, uses a wireless LAN-enabled network camera to collect footage from different events and locations, and the generation AI can integrate and edit them. For example, footage of a wedding reception and after-party can be integrated and edited. The network camera unit also uses a wireless LAN-enabled network camera to collect footage from different events and locations, and the generation AI can integrate and edit them. In this way, by using a wireless LAN-enabled network camera to collect footage from different events and locations, and the generation AI can integrate and edit them, it is possible to create a work that is full of realism.

[0047] The network camera unit uses a wireless LAN-enabled network camera to collect video according to a specific time period or situation, and the generation AI can edit it based on that. The network camera unit, for example, uses a wireless LAN-enabled network camera to collect video according to a specific time period or situation, and the generation AI can edit it based on that. For example, video from the start to the end of a wedding reception is collected and edited. The network camera unit also uses a wireless LAN-enabled network camera to collect video according to a specific time period or situation, and the generation AI can edit it based on that. In this way, a highly immersive work can be created by using a wireless LAN-enabled network camera to collect video according to a specific time period or situation, and the generation AI can edit it based on that.

[0048] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.

[0049] The video generation system also includes a voice synthesis unit. The voice synthesis unit can automatically generate appropriate narration and sound effects for important scenes extracted by the video analysis unit and incorporate them into the video. For example, in a video of a wedding reception, narration that enhances the emotion of the vows or moving speeches can be added. The voice synthesis unit can also generate sound effects in accordance with movements in the video, enhancing the sense of realism. This allows the video generation system to create moving works that appeal not only to the eyes but also to the ears.

[0050] The video generation system further includes a scene classification unit. The scene classification unit can classify the video material analyzed by the video analysis unit by scene type. For example, the scene classification unit can classify video of a wedding reception into speech scenes, dance scenes, cake-cutting scenes, etc. The scene classification unit can also evaluate the importance of each scene and prioritize editing of important scenes. This allows the video generation system to generate visually consistent works.

[0051] The video generation system further includes a subtitle generation unit. The subtitle generation unit can automatically generate subtitles based on the audio data analyzed by the video analysis unit and incorporate them into the video. For example, subtitles can be generated in real time for speeches and vows at a wedding reception. The subtitle generation unit also supports multiple languages ​​and can simultaneously display subtitles in different languages. This allows the video generation system to visually supplement information and accommodate a wider audience.

[0052] The video generation system further includes a scene prediction unit. The scene prediction unit can analyze past video data and predict the occurrence of future events and scenes. For example, in a video of a wedding reception, the next important scene to occur can be predicted based on past data, and that scene can be given priority in editing. The scene prediction unit can also prepare appropriate effects and music for the predicted scene in advance and incorporate them into the editing in real time. This allows the video generation system to achieve smooth editing based on predictions.

[0053] The video generation system further includes a scene connection unit. The scene connection unit can smoothly connect different scenes to build a consistent storyline. For example, in a video of a wedding reception, the transition from a speech scene to a dance scene can be made smooth. The scene connection unit can also automatically generate transition effects between scenes to achieve a visually natural flow. This allows the video generation system to generate a visually consistent work.

[0054] The video production system further includes a scene search unit. The scene search unit searches for relevant scenes from video footage based on keywords or phrases specified by the user and can incorporate them into the edit. For example, in a video of a wedding reception, the scene search unit searches for scenes based on keywords such as "vows" or "first dance." The scene search unit can also prioritize editing important scenes based on the search results. This allows the video production system to generate a customized work that meets the user's needs.

[0055] The processing flow of the first embodiment will be briefly explained below.

[0056] Step 1: The video analysis unit analyzes the video material. For example, the generation AI analyzes the video material and instantly identifies important scenes and highlights. The generation AI also analyzes the movements and gestures in the video material and extracts scenes with large movements as highlights. Furthermore, the generation AI uses voice recognition to detect important statements and audio events and extract them as highlights. Step 2: The editing department edits in real time based on the results of the analysis by the video analysis department. For example, the generation AI edits in real time based on the analyzed highlights to generate a complete work. The generation AI also automatically cuts and edits the video to match the rhythm and tempo of the music. Furthermore, the generation AI automatically adjusts the color tone and brightness within the video to generate a visually consistent work. Step 3: The facial recognition unit reads a group photo and performs facial recognition. For example, the generation AI reads a group photo and identifies guests using facial recognition technology. The generation AI also analyzes the facial expressions and movements of the guests and prioritizes editing scenes that show specific emotions. The generation AI then compares the footage with past video data of the guests and generates scenes that highlight the growth and changes of specific individuals. Step 4: The data reading unit reads the video data taken by the attendees. For example, the generation AI reads the video data taken by the attendees and uses it as video material. When analyzing the attendees' video data, the generation AI automatically stabilizes the video and removes noise to generate high-quality video. Furthermore, when analyzing the attendees' video data, the generation AI analyzes the audio in the video, extracts important remarks and audio events, and incorporates them into the editing. Step 5: The network camera unit collects video using a wireless LAN-enabled network camera. For example, the generation AI uses a wireless LAN-enabled network camera to stream video in real time and instantly analyzes and edits the video. The generation AI also uses a wireless LAN-enabled network camera to simultaneously collect video from multiple cameras, integrating and editing the video. The generation AI also uses a wireless LAN-enabled network camera to collect video from different events and locations, integrating and editing the video.

[0057] (Example 2) The video generation system according to the embodiment of the present invention is a system that uses AI technology to instantly analyze video material, edit it in real time, and generate a complete work. This allows the video generation system to analyze video material, edit it in real time, and generate an impressive work.

[0058] A video generation system according to an embodiment includes a video analysis unit, an editing unit, a facial recognition unit, a data reading unit, and a network camera unit. The video analysis unit analyzes video footage. For example, the generation AI analyzes the video footage and instantly identifies important scenes and highlights. The generation AI also analyzes the movements and gestures in the video footage and extracts scenes with significant movement as highlights. The generation AI also uses voice recognition to detect important statements and audio events and extract them as highlights. The editing unit performs real-time editing based on the results of the analysis by the video analysis unit. For example, the generation AI performs real-time editing based on the analyzed highlights to generate a complete work. The generation AI also automatically cuts and edits video to match the rhythm and tempo of the music. The generation AI also automatically adjusts the color tone and brightness in the video to generate a visually consistent work. The facial recognition unit reads a group photo and performs facial recognition. For example, the generation AI reads a group photo and identifies guests using facial recognition technology. The generation AI also analyzes the facial expressions and movements of guests and prioritizes editing scenes that highlight specific emotions. The generation AI also compares the guest's past video data to generate scenes that highlight the guest's growth and changes. The data import unit imports video data shot by attendees. For example, the generation AI imports video data shot by attendees and uses it as video material. When analyzing attendees' video data, the generation AI automatically stabilizes the video and removes noise to generate high-quality video. When analyzing attendees' video data, the generation AI also analyzes the audio in the video, extracting important comments and audio events and incorporating them into the editing. The network camera unit collects video using a wireless LAN-enabled network camera. For example, the generation AI uses a wireless LAN-enabled network camera to stream video in real time and instantly analyze and edit the video. The generation AI also uses a wireless LAN-enabled network camera to simultaneously collect video from multiple cameras and integrate and edit the video. The generative AI also uses wireless LAN-enabled network cameras to collect footage from different events and locations, then integrates and edits it.As a result, the video generation system according to the embodiment can analyze video footage, edit it in real time, and create a moving work. For example, the system can analyze video footage of a wedding reception, extract important scenes, and edit it in real time to create a moving work. It can also read video data taken by attendees and create a realistic video. It can also use a wireless LAN-compatible network camera to collect video in real time and instantly analyze and edit it.

[0059] The video analysis unit can use voice recognition to detect important statements and audio events and extract them as highlights. For example, when the generation AI analyzes video material, the video analysis unit uses voice recognition technology to detect important statements and audio events. For example, scenes with important audio content, such as wedding reception speeches and vows, can be automatically extracted. The generation AI also uses voice recognition technology to analyze audio within the video, extract important statements and audio events, and incorporate them into the editing. This makes it possible to highlight important scenes by detecting important statements and audio events using voice recognition and extracting them as highlights.

[0060] The video analysis unit analyzes movements and gestures within the video and can extract scenes with large movements as highlights. For example, when the generation AI analyzes video material, the video analysis unit analyzes movements within the video and extracts scenes with large movements as highlights. For example, scenes with a lot of movement, such as dance scenes or the moment of cutting a cake, are automatically extracted. The video analysis unit also allows the generation AI to analyze gestures within the video and extract scenes with large movements as highlights. This allows scenes with a lot of movement to be emphasized by analyzing movements and gestures within the video and extracting scenes with a lot of movement as highlights.

[0061] The video analysis unit can use the emotion estimation function to estimate emotions from the facial expressions and tone of voice of people in the video and extract emotionally significant scenes as highlights. The video analysis unit, for example, uses the emotion estimation function to analyze the facial expressions of people in the video and extract emotionally significant scenes as highlights. For example, scenes of smiling or crying are automatically detected and extracted. The video analysis unit also uses the emotion estimation function to analyze the tone of voice of people in the video and extract emotionally significant scenes as highlights. For example, scenes with moving speeches or vows are extracted. In this way, emotional scenes can be emphasized by using the emotion estimation function to estimate emotions from the facial expressions and tone of voice of people in the video and extract emotionally significant scenes as highlights.

[0062] The video analysis unit can simultaneously analyze video from different camera angles and viewpoints to generate highlights from multiple viewpoints. For example, the video analysis unit uses a generation AI to simultaneously analyze video from different camera angles to generate highlights from multiple viewpoints. For example, the video analysis unit analyzes video from multiple cameras at a wedding reception to extract important scenes from multiple angles. The video analysis unit also uses a generation AI to analyze video from different viewpoints to generate highlights from multiple viewpoints. This allows the video analysis unit to simultaneously analyze video from different camera angles and viewpoints to generate highlights from multiple viewpoints, thereby providing a more realistic video.

[0063] The video analysis unit can compare it with past event data, detect similar scenes and patterns, and extract them as highlights. For example, when the generation AI analyzes video material, the video analysis unit compares it with past event data, detects similar scenes, and extracts them as highlights. For example, it compares it with footage of a past wedding reception to extract similar important scenes. The video analysis unit also allows the generation AI to compare it with past event data, detect similar patterns, and extract them as highlights. This makes it possible to highlight important scenes by comparing it with past event data, detecting similar scenes and patterns, and extracting them as highlights.

[0064] The video analysis unit can prioritize analyzing scenes in which the user has a specific emotion and generate highlights based on that emotion. The video analysis unit, for example, uses an emotion estimation function to prioritize analyzing scenes in which the user has a specific emotion and generate highlights based on that emotion. For example, it prioritizes extracting scenes in which the user was moved. The video analysis unit also uses a generation AI to analyze the user's emotion and generate highlights based on that emotion. This allows emotional scenes to be emphasized by prioritized analysis of scenes in which the user has a specific emotion and generating highlights based on that emotion.

[0065] The editing department can automatically cut and edit video to match the rhythm and tempo of the music. For example, the generative AI automatically cuts and edits video to match the rhythm and tempo of the music during real-time editing. For example, video of a wedding reception can be cut to match the beat of the music to generate a rhythmic video. The editing department can also automatically cut and edit video to match the rhythm and tempo of the music using the generative AI. This allows for harmony between the visual and auditory senses by automatically cutting and editing video to match the rhythm and tempo of the music.

[0066] The editing department can automatically adjust the color tone and brightness within the video to generate a visually consistent work. For example, the generative AI can automatically adjust the color tone and brightness within the video during real-time editing to generate a visually consistent work. For example, the generative AI can unify the color tones of footage of a wedding reception to create a professional finish. The editing department can also have the generative AI automatically adjust the color tone and brightness within the video to generate a visually consistent work. This automatically adjusts the color tone and brightness within the video to generate a visually consistent work, allowing for a professional finish.

[0067] The editorial department can use the emotion estimation function to automatically select music and effects that match the user's emotions and generate an emotional work. The editorial department, for example, uses the emotion estimation function to automatically select music that matches the user's emotions and generate an emotional work. For example, emotional music is selected for an emotional scene. The editorial department also uses the emotion estimation function to automatically select effects that match the user's emotions and generate an emotional work. In this way, by using the emotion estimation function to automatically select music and effects that match the user's emotions and generate an emotional work, it is possible to achieve harmony between the visual and auditory senses.

[0068] During real-time editing, the editing department can edit the video based on keywords and phrases specified by the user to emphasize a particular message. For example, during real-time editing, the editing department can edit the video based on keywords and phrases specified by the user to emphasize a particular message. For example, editing can be done based on keywords such as "love" or "gratitude." Furthermore, during real-time editing, the generative AI can edit the video based on keywords and phrases specified by the user to emphasize a particular message. In this way, by editing the video based on keywords and phrases specified by the user during real-time editing to emphasize a particular message, it is possible to visually convey a message.

[0069] The editorial department can use the emotion estimation function to prioritize editing scenes that move the user most and compose a work around those scenes. The editorial department, for example, can use the emotion estimation function to prioritize editing scenes that move the user most and compose a work around those scenes. For example, editing can be done around scenes of moving speeches or vow kisses. The editorial department can also have the generation AI use the emotion estimation function to prioritize editing scenes that move the user most and compose a work around those scenes. In this way, by using the emotion estimation function to prioritize editing scenes that move the user most and compose a work around those scenes, a moving work can be generated.

[0070] The facial recognition unit analyzes the facial expressions and movements of guests and can prioritize editing of scenes that show particular emotions. For example, when the generation AI recognizes a guest's face, the facial recognition unit analyzes the guest's facial expressions and prioritizes editing of scenes that show particular emotions. For example, scenes of smiling or crying are automatically detected and edited. The facial recognition unit also allows the generation AI to analyze the guest's movements and prioritize editing of scenes that show particular emotions. This allows emotional scenes to be emphasized by analyzing the guest's facial expressions and movements and prioritizing editing of scenes that show particular emotions.

[0071] The facial recognition unit can compare the video with past video data of guests and generate scenes that highlight the growth and changes of specific individuals. For example, when the generation AI recognizes a face, the facial recognition unit compares the video with past video data of guests and generates scenes that highlight the growth and changes of specific individuals. For example, the generation AI compares video footage of past weddings with the current wedding and edits it. The facial recognition unit also allows the generation AI to compare the video with past video data of guests and generate scenes that highlight the growth and changes of specific individuals. This allows the generation AI to compare the video with past video data of guests and generate scenes that highlight the growth and changes of specific individuals, creating a moving work.

[0072] The facial recognition unit can use the emotion estimation function to estimate the emotions of guests and prioritize editing of emotionally important scenes. The facial recognition unit, for example, uses the emotion estimation function to estimate the emotions of guests and prioritize editing of emotionally important scenes. For example, it prioritizes editing of moving speeches and vow kiss scenes. Furthermore, the facial recognition unit allows the generation AI to use the emotion estimation function to estimate the emotions of guests and prioritize editing of emotionally important scenes. In this way, by using the emotion estimation function to estimate the emotions of guests and prioritize editing of emotionally important scenes, it is possible to generate a moving work.

[0073] The facial recognition unit can analyze the clothing and accessories of guests and generate video that matches a specific theme or style. For example, when the generation AI recognizes their faces, the facial recognition unit analyzes the clothing and accessories of guests and generates video that matches a specific theme or style. For example, it generates video that matches the dress code of a wedding reception. The facial recognition unit also allows the generation AI to analyze the clothing and accessories of guests and generate video that matches a specific theme or style. This makes it possible to generate visually consistent works by analyzing the clothing and accessories of guests and generating video that matches a specific theme or style.

[0074] The facial recognition unit can analyze the relationships between guests and prioritize editing scenes based on those relationships. For example, when the generation AI recognizes their faces, the facial recognition unit analyzes the relationships between guests and prioritizes editing scenes based on those relationships. For example, scenes with family and friends are prioritized. The facial recognition unit also allows the generation AI to analyze the relationships between guests and prioritize editing scenes based on those relationships. This makes it possible to create moving works by analyzing the relationships between guests and prioritize editing scenes based on those relationships.

[0075] The facial recognition unit can use the emotion estimation function to prioritize editing scenes that will move the guests most and compose a work around those scenes. The facial recognition unit can, for example, use the emotion estimation function to prioritize editing scenes that will move the guests most and compose a work around those scenes. For example, editing scenes of moving speeches or vow kisses can be edited around them. The facial recognition unit can also use the emotion estimation function to prioritize editing scenes that will move the guests most and compose a work around those scenes. In this way, by using the emotion estimation function to prioritize editing scenes that will move the guests most and compose a work around those scenes, a moving work can be generated.

[0076] The data reading unit can automatically stabilize the video and remove noise when analyzing the attendees' video data, thereby generating high-quality video. For example, the data reading unit can automatically stabilize the video and generate high-quality video when the generation AI analyzes the attendees' video data. For example, the data reading unit can perform image stabilization to generate stable video. Furthermore, the data reading unit can automatically remove noise when the generation AI analyzes the attendees' video data, thereby generating high-quality video. As a result, the data reading unit can automatically stabilize the video and remove noise when analyzing the attendees' video data, thereby generating high-quality video, thereby achieving a professional finish.

[0077] When analyzing attendees' video data, the data loading unit analyzes the audio in the video, extracts important remarks and audio events, and incorporates them into the editing. For example, when the generation AI analyzes attendees' video data, the data loading unit analyzes the audio in the video, extracts important remarks and audio events, and incorporates them into the editing. For example, speeches and moving words are extracted and edited. The data loading unit also enables the generation AI to analyze the audio in the video, extract important remarks and audio events, and incorporate them into the editing. This allows important scenes to be emphasized by analyzing the audio in the video, extracting important remarks and audio events, and incorporating them into the editing.

[0078] The data reading unit can use the emotion estimation function to estimate the emotions of the attendees and prioritize editing of emotionally important scenes. The data reading unit, for example, uses the emotion estimation function to estimate the emotions of the attendees and prioritize editing of emotionally important scenes. For example, it prioritizes editing of scenes of moving speeches and vow kisses. The data reading unit also uses the emotion estimation function of the generation AI to estimate the emotions of the attendees and prioritize editing of emotionally important scenes. In this way, by using the emotion estimation function to estimate the emotions of the attendees and prioritize editing of emotionally important scenes, it is possible to generate a moving work.

[0079] When analyzing the video data of attendees, the data loading unit can combine footage from different viewpoints and angles to generate highly immersive video. For example, when the generation AI analyzes the video data of attendees, the data loading unit can combine footage from different viewpoints and angles to generate highly immersive video. For example, the data loading unit can combine and edit footage shot by multiple attendees. In addition, the data loading unit can allow the generation AI to combine footage from different viewpoints and angles to generate highly immersive video. As a result, when analyzing the video data of attendees, the data loading unit can combine footage from different viewpoints and angles to generate highly immersive video, making it possible to create visually appealing works.

[0080] The data loading unit can edit the video based on a specific theme or storyline when analyzing the video data of attendees, thereby generating a consistent work. For example, when the generation AI analyzes the video data of attendees, the data loading unit can edit the video based on a specific theme or storyline to generate a consistent work. For example, editing based on a wedding reception theme. The data loading unit can also edit the video based on a specific theme or storyline when the generation AI analyzes the video data of attendees, thereby generating a consistent work. In this way, when analyzing the video data of attendees, the data loading unit can edit the video based on a specific theme or storyline to generate a consistent work, thereby generating a visually appealing work.

[0081] The network camera unit uses a wireless LAN-enabled network camera to stream video in real time, and the generation AI can instantly analyze and edit that video. The network camera unit, for example, uses a wireless LAN-enabled network camera to stream video in real time, and the generation AI can instantly analyze and edit that video. For example, video of a wedding reception can be analyzed in real time to extract important scenes. The network camera unit also uses a wireless LAN-enabled network camera to stream video in real time, and the generation AI can instantly analyze and edit that video. This makes it possible to stream video in real time using a wireless LAN-enabled network camera, and the generation AI can instantly analyze and edit that video, creating impressive works in real time.

[0082] The network camera unit uses wireless LAN-enabled network cameras to simultaneously collect footage from multiple cameras, and the generation AI can integrate and edit it. The network camera unit, for example, uses wireless LAN-enabled network cameras to simultaneously collect footage from multiple cameras, and the generation AI can integrate and edit it. For example, footage from multiple cameras of a wedding reception can be integrated and edited. The network camera unit also uses wireless LAN-enabled network cameras to simultaneously collect footage from multiple cameras, and the generation AI can integrate and edit it. This allows for the creation of a work that is full of realism.

[0083] The network camera unit can use the emotion estimation function to analyze the emotions of people in the video captured by the network camera and prioritize editing of emotionally significant scenes. For example, the network camera unit can use the emotion estimation function to analyze the emotions of people in the video captured by the network camera and prioritize editing of emotionally significant scenes. For example, it can prioritize editing of scenes with moving speeches or vow kisses. In addition, the network camera unit's generation AI uses the emotion estimation function to analyze the emotions of people in the video captured by the network camera and prioritize editing of emotionally significant scenes. In this way, by using the emotion estimation function to analyze the emotions of people in the video captured by the network camera and prioritize editing of emotionally significant scenes, it is possible to create a moving work.

[0084] The network camera unit uses a wireless LAN-enabled network camera to collect footage from different events and locations, and the generation AI can integrate and edit them. The network camera unit, for example, uses a wireless LAN-enabled network camera to collect footage from different events and locations, and the generation AI can integrate and edit them. For example, footage of a wedding reception and after-party can be integrated and edited. The network camera unit also uses a wireless LAN-enabled network camera to collect footage from different events and locations, and the generation AI can integrate and edit them. In this way, by using a wireless LAN-enabled network camera to collect footage from different events and locations, and the generation AI can integrate and edit them, it is possible to create a work that is full of realism.

[0085] The network camera unit uses a wireless LAN-enabled network camera to collect video according to a specific time period or situation, and the generation AI can edit it based on that. The network camera unit, for example, uses a wireless LAN-enabled network camera to collect video according to a specific time period or situation, and the generation AI can edit it based on that. For example, video from the start to the end of a wedding reception is collected and edited. The network camera unit also uses a wireless LAN-enabled network camera to collect video according to a specific time period or situation, and the generation AI can edit it based on that. In this way, a highly immersive work can be created by using a wireless LAN-enabled network camera to collect video according to a specific time period or situation, and the generation AI can edit it based on that.

[0086] The network camera unit uses the emotion estimation function to analyze the emotions of people in the video captured by the network camera in real time and prioritize editing of emotionally significant scenes. The network camera unit, for example, uses the emotion estimation function to analyze the emotions of people in the video captured by the network camera in real time and prioritize editing of emotionally significant scenes. For example, it prioritizes editing of scenes with moving speeches or vow kisses. In addition, the network camera unit's generation AI uses the emotion estimation function to analyze the emotions of people in the video captured by the network camera in real time and prioritize editing of emotionally significant scenes. In this way, by using the emotion estimation function to analyze the emotions of people in the video captured by the network camera in real time and prioritize editing of emotionally significant scenes, it is possible to create a moving work.

[0087] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.

[0088] The video generation system also includes a voice synthesis unit. The voice synthesis unit can automatically generate appropriate narration and sound effects for important scenes extracted by the video analysis unit and incorporate them into the video. For example, in a video of a wedding reception, narration that enhances the emotion of the vows or moving speeches can be added. The voice synthesis unit can also generate sound effects in accordance with movements in the video, enhancing the sense of realism. This allows the video generation system to create moving works that appeal not only to the eyes but also to the ears.

[0089] The video generation system further includes a scene classification unit. The scene classification unit can classify the video material analyzed by the video analysis unit by scene type. For example, the scene classification unit can classify video of a wedding reception into speech scenes, dance scenes, cake-cutting scenes, etc. The scene classification unit can also evaluate the importance of each scene and prioritize editing of important scenes. This allows the video generation system to generate visually consistent works.

[0090] The video generation system further includes a background music selection unit. The background music selection unit can automatically select appropriate background music for the video material analyzed by the video analysis unit and incorporate it into the video. For example, in a video of a wedding reception, it can select moving music for moving scenes and happy music for happy scenes. The background music selection unit can also adjust the music to match the tempo and rhythm of the video, achieving harmony between the visual and auditory senses. This allows the video generation system to generate works that are moving both visually and aurally.

[0091] The video generation system further includes a subtitle generation unit. The subtitle generation unit can automatically generate subtitles based on the audio data analyzed by the video analysis unit and incorporate them into the video. For example, subtitles can be generated in real time for speeches and vows at a wedding reception. The subtitle generation unit also supports multiple languages ​​and can simultaneously display subtitles in different languages. This allows the video generation system to visually supplement information and accommodate a wider audience.

[0092] The video generation system further includes an emotion feedback unit. The emotion feedback unit can collect emotions felt by the user while watching in real time and adjust the editing of the video based on the feedback. For example, the emotion feedback unit can add emotional music or effects to emphasize scenes that move the user. The emotion feedback unit can also adjust the tempo and rhythm of the video based on the user's emotions to achieve harmony between the visual and auditory senses. This allows the video generation system to generate a customized work that matches the user's emotions.

[0093] The video generation system further includes a scene prediction unit. The scene prediction unit can analyze past video data and predict the occurrence of future events and scenes. For example, in a video of a wedding reception, the next important scene to occur can be predicted based on past data, and that scene can be given priority in editing. The scene prediction unit can also prepare appropriate effects and music for the predicted scene in advance and incorporate them into the editing in real time. This allows the video generation system to achieve smooth editing based on predictions.

[0094] The video generation system further includes an emotion analysis unit. The emotion analysis unit performs detailed analysis of emotional changes in the video material analyzed by the video analysis unit and can perform editing based on the flow of emotions. For example, in a video of a wedding reception, the emotion analysis unit can smoothly transition from moving scenes to happy scenes. The emotion analysis unit can also add appropriate effects and music to emphasize emotional peaks. This allows the video generation system to generate moving works that emphasize the flow of emotions.

[0095] The video generation system further includes a scene connection unit. The scene connection unit can smoothly connect different scenes to build a consistent storyline. For example, in a video of a wedding reception, the transition from a speech scene to a dance scene can be made smooth. The scene connection unit can also automatically generate transition effects between scenes to achieve a visually natural flow. This allows the video generation system to generate a visually consistent work.

[0096] The video generation system further includes an emotion emphasis unit. The emotion emphasis unit can add effects and music to emphasize emotions in emotionally significant scenes analyzed by the video analysis unit. For example, effects and music that enhance the emotion can be added to a moving speech scene at a wedding reception. The emotion emphasis unit can also adjust the color tone and brightness of the video to emphasize the peak of emotion. This allows the video generation system to generate emotionally emphasized and moving works.

[0097] The video production system further includes a scene search unit. The scene search unit searches for relevant scenes from video footage based on keywords or phrases specified by the user and can incorporate them into the edit. For example, in a video of a wedding reception, the scene search unit searches for scenes based on keywords such as "vows" or "first dance." The scene search unit can also prioritize editing important scenes based on the search results. This allows the video production system to generate a customized work that meets the user's needs.

[0098] The processing flow of the second embodiment will be briefly explained below.

[0099] Step 1: The video analysis unit analyzes the video material. For example, the generation AI analyzes the video material and instantly identifies important scenes and highlights. The generation AI also analyzes the movements and gestures in the video material and extracts scenes with large movements as highlights. Furthermore, the generation AI uses voice recognition to detect important statements and audio events and extract them as highlights. Step 2: The editing department edits in real time based on the results of the analysis by the video analysis department. For example, the generation AI edits in real time based on the analyzed highlights to generate a complete work. The generation AI also automatically cuts and edits the video to match the rhythm and tempo of the music. Furthermore, the generation AI automatically adjusts the color tone and brightness within the video to generate a visually consistent work. Step 3: The facial recognition unit reads a group photo and performs facial recognition. For example, the generation AI reads a group photo and identifies guests using facial recognition technology. The generation AI also analyzes the facial expressions and movements of the guests and prioritizes editing scenes that show specific emotions. The generation AI then compares the footage with past video data of the guests and generates scenes that highlight the growth and changes of specific individuals. Step 4: The data reading unit reads the video data taken by the attendees. For example, the generation AI reads the video data taken by the attendees and uses it as video material. When analyzing the attendees' video data, the generation AI automatically stabilizes the video and removes noise to generate high-quality video. Furthermore, when analyzing the attendees' video data, the generation AI analyzes the audio in the video, extracts important remarks and audio events, and incorporates them into the editing. Step 5: The network camera unit collects video using a wireless LAN-enabled network camera. For example, the generation AI uses a wireless LAN-enabled network camera to stream video in real time and instantly analyzes and edits the video. The generation AI also uses a wireless LAN-enabled network camera to simultaneously collect video from multiple cameras, integrating and editing the video. The generation AI also uses a wireless LAN-enabled network camera to collect video from different events and locations, integrating and editing the video.

[0100] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0101] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> Examples of generative AIs include the data generation model 58, such as a neural network model (e.g., a neural network model), and a neural network model (e.g., a neural network model). The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating speech, text data indicating text, and image data indicating an image is also input to the data generation model 58. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specification processing unit 290 performs the above-mentioned specification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0102] Furthermore, the processing by the data processing system 10 described above is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0103] [Second embodiment] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0104] 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0105] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.

[0106] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0107] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.

[0108] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0109] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0110] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0111] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0112] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0113] In the smart glasses 214, the specific processing is performed by the processor 46. A specific processing program 60 is stored in the storage 50. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. Note that the smart glasses 214 may have a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59.

[0114] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.

[0115] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0116] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0117] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or an external device, etc., and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or an external device, etc.

[0118] [Third embodiment] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0119] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0120] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.

[0121] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0122] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.

[0123] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0124] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0125] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0126] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0127] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0128] In the headset type terminal 314, the identification process is performed by the processor 46. A identification program 60 is stored in the storage 50. The processor 46 reads the identification program 60 from the storage 50 and executes the read identification program 60 on the RAM 48. The identification process is realized by the processor 46 operating as a control unit 46A in accordance with the identification program 60 executed on the RAM 48. Note that the headset type terminal 314 may also have a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59.

[0129] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.

[0130] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0131] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0132] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset type terminal 314, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset type terminal 314. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the headset type terminal 314 or an external device, etc., and the headset type terminal 314 acquires or collects information required for processing from the data processing device 12 or an external device, etc.

[0133] [Fourth embodiment] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0134] 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0135] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.

[0136] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0137] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.

[0138] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS image sensor or a CCD image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0139] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0140] The control object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0141] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0142] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0143] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0144] In the robot 414, the processor 46 performs the identification process. A identification program 60 is stored in the storage 50. The processor 46 reads the identification program 60 from the storage 50 and executes the read identification program 60 on the RAM 48. The identification process is realized by the processor 46 operating as a control unit 46A in accordance with the identification program 60 executed on the RAM 48. The robot 414 may have a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59.

[0145] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.

[0146] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0147] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0148] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or an external device, etc., and the robot 414 acquires or collects information required for processing from the data processing device 12 or an external device, etc.

[0149] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0150] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion encompasses both emotions and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[0151] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[0152] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[0153] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. Emotions can also be created for robots, cars, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on speech emotion recognition and brain physiological signal analysis systems for emotions, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the area called "reaction," where sensation is dominant. The right half of the emotion map lists emotions belonging to the area called "situation," where situational awareness is dominant.

[0154] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[0155] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[0156] In the above embodiment, an example was given in which a specific process is performed by one computer 22, but the technology disclosed herein is not limited to this, and distributed processing of the specific process may be performed by multiple computers including computer 22.

[0157] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[0158] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0159] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[0160] The hardware resource for executing a specific process can be any of the following processors: A CPU is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. A dedicated electrical circuit, such as a field-programmable gate array (FPGA), a programmable logic device (PLD), or an application-specific integrated circuit (ASIC), is a processor with a circuit configuration specifically designed to execute a specific process. Each processor has built-in or connected memory, and uses the memory to execute the specific process.

[0161] The hardware resource that executes the specific process may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.

[0162] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[0163] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[0164] In the above example, the first to fourth embodiments have been described separately, but some or all of these embodiments may be combined. The smart device 14, smart glasses 214, headset terminal 314, and robot 414 are merely examples, and they may be combined, or other devices may be used. In the above example, the first and second embodiments have been described separately, but they may be combined.

[0165] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[0166] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference. [Explanation of symbols]

[0167] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot

Claims

1. a video analysis unit that analyzes video material; an editing unit that performs editing in real time based on the results of analysis by the video analysis unit; A face recognition unit that reads a group photo and performs face recognition; a data reading unit that reads video data taken by attendees; A network camera unit that collects video using a wireless LAN-compatible network camera. A system characterized by:

2. The video analysis unit Uses speech recognition to detect important utterances and audio events and extract them as highlights 2. The system of claim 1.

3. The video analysis unit Simultaneously analyze footage from different camera angles and viewpoints to generate highlights from multiple perspectives 2. The system of claim 1.

4. The editorial department Automatically cut and edit the video in accordance with the rhythm and tempo of the music 2. The system of claim 1.

5. The face authentication unit Analyzing the facial expressions and movements of guests and prioritizing editing scenes that convey specific emotions 2. The system of claim 1.

6. The data reading unit When analyzing the video data of the attendees, the video is automatically stabilized and noise is removed to generate high-quality video.

2. The system of claim 1.

7. The network camera unit Using the wireless LAN-enabled network camera, the video is streamed in real time, and the generation AI instantly analyzes and edits the video.

2. The system of claim 1.

8. The video analysis unit Using emotion estimation functionality, emotions are estimated from facial expressions and tone of voice of people in the video, and emotionally significant scenes are extracted as highlights.

2. The system of claim 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A