system
A system with real-time video and audio collection, server analysis, and attendee integration addresses the challenge of skilled cameramen shortages by automatically editing event highlights, ensuring high-quality video production.
Patent Information
- Application Number
- JP2024122752
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-02-10
AI Technical Summary
Securing skilled cameramen for events like wedding receptions is difficult, and incorporating attendee footage is challenging, necessitating a system for real-time extraction and editing of highlight scenes.
A system that uses multiple cameras to collect video and audio in real-time, a server for analysis and editing, and attendee uploads to integrate footage, utilizing facial recognition and AI for scene identification and real-time editing.
Enables high-quality video recording and editing without relying on skilled personnel, meeting customer demands by automatically extracting and integrating footage into a final video.
Smart Images

Figure 2026021070000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] At events such as wedding receptions, highly skilled cameramen are required to capture the perfect moments on video, but securing such personnel is becoming increasingly difficult, making it difficult to meet customer demands. Furthermore, incorporating footage shot by attendees themselves to create more realistic footage is also a challenge. To address these issues, a system is needed that can automatically extract and edit the best scenes in real time. [Means for solving the problem]
[0005] This invention provides a means for collecting video and audio from multiple cameras in real time and transmitting them to a server via wireless LAN. The server uses facial recognition to identify guests, analyzes the video and audio using AI, and identifies highlight scenes. The server performs real-time editing based on the identified highlight scenes. Furthermore, attendees can upload footage they have taken to the server using their devices, and the server integrates the footage into existing highlight footage. All video data is used to generate a final video, and re-edits the video based on re-editing requests. By outputting the generated final video and re-edited video, a method is provided for preserving high-quality video records while meeting customer requests.
[0006] A "camera" is a device that collects video and audio and transmits the data via wireless LAN.
[0007] "Wireless LAN" is a technology for connecting to a network via wireless communication.
[0008] The "server" is a remote control device that analyzes and edits video and audio data received from multiple cameras.
[0009] "Facial recognition" is a technology that identifies people's faces from images such as group photos.
[0010] "AI" is a system that uses artificial intelligence technology to analyze data and extract specific patterns and characteristics.
[0011] A "highlight scene" is a scene that refers to an important or moving moment in an event.
[0012] A "terminal" is a device that is operated by a user and that transmits and receives data.
[0013] "Upload" is the operation of transferring data from a local device to a server over a network.
[0014] "Editing" is the process of processing collected video and audio data, cutting out unnecessary parts, and emphasizing specific scenes.
[0015] "Integration" is the process of combining multiple pieces of data into one format.
[0016] "Re-editing" refers to the process of adding new processing to video data that has already been edited.
[0017] "Output" refers to the operation of providing the video data generated by the server in a format that is accessible to the user. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] This invention is a system that uses multiple cameras to collect video and audio in real time, and a server analyzes and edits the video and audio. The specific program processing is explained below in natural language.
[0040] Camera Initial Settings
[0041] The terminal installs wireless LAN-compatible cameras in various locations in the venue. The cameras are placed on the ceiling, on tabletops, and around tables, and the terminal registers each camera's IP address, field of view, and installation location on the server.
[0042] Video collection and transmission
[0043] Cameras collect video and audio data from the entire venue in real time, and this data is sent to a server via wireless LAN.
[0044] Video and audio analysis
[0045] The server receives and analyzes the streaming data in real time. First, it uses facial recognition to identify guests based on group photos and video clips. Next, it uses AI to analyze the video and audio data and identify highlights based on emotional expressions and volume changes.
[0046] Video editing
[0047] The server then edits the identified highlights in real time, prioritizing important scenes such as the bride and groom's entrance, speeches, toasts, and cake cutting. The edited footage is then temporarily saved.
[0048] Attendee Video Integration
[0049] Users (attendees) upload video data taken with their smartphones or cameras to the device. The device then sends the uploaded data to the server, which receives and analyzes it. The received video data is integrated with existing highlight footage to create a more realistic video.
[0050] Generating the final video
[0051] The server generates the final video from all the video data, resulting in a complete video at the end of the event. The generated video file is then provided to users in an easily accessible format.
[0052] Re-editing function
[0053] The user sends a request to the server via their device to re-edit a specific scene. For example, if the user has a specific request, such as "I want more emphasis on the birthday cake," the server re-analyzes the video data and makes a new edit. The re-edited video is then provided to the user.
[0054] Specific examples
[0055] 1. Video recording of a wedding reception
[0056] The terminal places wireless LAN-compatible cameras at various locations throughout the venue and performs initial setup.
[0057] The camera captures the moment when the bride and groom enter, the guests give speeches, the toast, and the cake cutting in real time, and sends the video and audio data to the server.
[0058] The server analyzes the collected data and uses facial recognition to identify the bride and groom and important guests. Using AI, it extracts highlights based on emotional expressions and volume changes and edits them in real time.
[0059] Users (attendees) upload videos taken with their smartphones to the device and transfer them to the server.
[0060] The server then combines the uploaded footage with existing edited footage to create a richer video.
[0061] The server generates the final edited video and makes it ready to be served.
[0062] In this way, advanced video recording and editing can be achieved using AI and networks, without relying on the skills of cameramen or editors.
[0063] The processing flow will be explained below.
[0064] Step 1:
[0065] The terminal installs wireless LAN-compatible cameras in advance at various locations in the venue and registers information such as the camera's IP address, field of view, and installation location on the server.
[0066] Step 2:
[0067] The cameras begin collecting video and audio from the entire venue in real time and stream this data to a server via wireless LAN.
[0068] Step 3:
[0069] The server receives the streaming data sent from the camera in real time and begins analyzing the video and audio.
[0070] Step 4:
[0071] The server performs facial recognition based on group photos and data acquired in advance to identify guests.
[0072] Step 5:
[0073] The server uses AI to analyze video and audio, detecting emotional expressions and changes in volume to identify highlight scenes.
[0074] Step 6:
[0075] The server then prioritizes editing the identified highlight scenes, cutting out unnecessary parts to optimize scene transitions, and temporarily stores the edited footage.
[0076] Step 7:
[0077] Users (attendees) upload video data taken with their smartphones or cameras to the device.
[0078] Step 8:
[0079] The device sends the uploaded video data of attendees to the server, which receives and analyzes it.
[0080] Step 9:
[0081] The server then combines the footage uploaded by attendees with existing highlight footage to create a more immersive video.
[0082] Step 10:
[0083] The server generates the final video from all the video data and provides the video file in an easily accessible format for the user.
[0084] Step 11:
[0085] The user sends a request to the server via their device to re-edit a specific scene.
[0086] Step 12:
[0087] The server re-analyzes the video based on the re-editing request, performs a new edit, and provides the re-edited video to the user.
[0088] Example 1
[0089] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0090] Conventional video recording systems rely on manual processes for collecting, analyzing, and editing video and audio, resulting in limitations in real-time performance and accuracy. This makes efficient, high-quality video editing particularly difficult in situations where large amounts of data must be collected in a short period of time, such as events and meetings. It is also cumbersome for users to easily integrate and re-edit footage they have shot themselves. There is a demand for a system that can solve these issues and achieve efficient, high-precision video recording and editing.
[0091] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0092] In this invention, the server includes means for collecting video and audio from a camera in real time, means for transmitting the collected video and audio to the server via wireless communication, means for the server to identify guests using facial recognition, means for the server to analyze the video and audio using artificial intelligence and identify highlight scenes, means for editing in real time based on the highlight scenes identified by the server, means for uploading video data taken by attendees to the server using a terminal, means for integrating the uploaded video data of attendees with existing highlight videos, means for generating a final video using all the video data, means for re-editing based on a request to re-edit specific scenes, and means for outputting the generated final video and re-edited video, thereby enabling efficient and high-quality video recording and editing.
[0093] A "camera" is a photographic device that collects video and audio from within the venue in real time and transmits them to a server.
[0094] "Wireless communication" is a general term for technology that uses radio waves to send and receive data.
[0095] A "server" is a computer system that receives data from multiple cameras and devices via a network and analyzes, edits, saves, and outputs the data.
[0096] "Facial recognition" is a technology that analyzes faces contained in video data and identifies individuals.
[0097] "Artificial intelligence" is a data analysis technology that uses machine learning and deep learning to analyze video based on emotional expressions, volume changes, etc.
[0098] "Highlight scenes" refer to important moments or scenes with strong emotional expressions within video data.
[0099] "Real-time editing" refers to the process of collecting video data and simultaneously analyzing and editing it to create a format that can be output immediately.
[0100] An "attendee" is a person attending an event, meeting, etc., including anyone who captures video or audio on their own device.
[0101] "Terminals" are devices used by attendees to upload video data to the server. These devices include smartphones and tablets.
[0102] "Integration" refers to the process of analyzing multiple video data sets and editing and combining them into a series of flows.
[0103] The "final video" is the completed video edited based on all the collected video data.
[0104] "Re-editing" refers to the process of adding new edits to already edited video data based on a request from a user.
[0105] "Generation" refers to the process of creating the final video from the collected data.
[0106] "Output" refers to providing the generated Final Video or Re-Edited Video in a user-accessible format.
[0107] The present invention is a system that uses multiple cameras to collect video and audio in real time, and a server analyzes and edits the collected video and audio. Specifically, the system is implemented according to the following steps:
[0108] First, wireless communication-enabled cameras are installed in various locations throughout the venue. The cameras are placed on the ceiling, on tabletops, and around tables, and the IP address, field of view, and installation location of each camera are registered on the server by the terminal. At this time, a dedicated setting application is used to save the camera information in the server database.
[0109] Next, the camera collects video and audio from the entire venue in real time. The collected data is sent to the server via wireless communication as an H.264 video stream. The camera automatically starts transmitting data and provides a live feed to a URL specified by the server.
[0110] The server receives and analyzes the streaming data in real time. First, it performs facial recognition based on group photos and pre-collected video clips to identify guests. Specifically, the server calls an AI model to perform facial recognition and identify individual guests. Next, it uses the AI model to analyze the video and audio data and identify highlight scenes based on emotional expressions and volume changes. For example, it prioritizes the extraction of scenes featuring the bride and groom or scenes with moving speeches.
[0111] The server then edits the identified highlights in real time. Specifically, the server uses video editing software to edit the video. Important scenes, such as the bride and groom's entrance, the toast, and the cake cutting, are extracted as clips and placed on the timeline. The edited video is then temporarily saved in an appropriate format.
[0112] Users (attendees) upload video data they have taken with their own devices to their terminals. This data is then sent to the server using a dedicated upload application. The video data received by the server is synchronized and integrated with existing highlight footage. The server analyzes this data and inserts additional footage at the appropriate times to create a more immersive final video.
[0113] The server then combines all the video data to create the final video. Specifically, it combines each edited clip and organizes the overall timeline. The video format is a common format such as MP4 or MOV. The final video file is then uploaded to a cloud storage service, and a shared link is provided for users to easily access.
[0114] It is also possible for users to send requests to the server via their device to re-edit specific scenes. For example, if a user has a specific request, such as "I want more emphasis on the toast scene," the user enters the details in a dedicated form and sends it to the server. The server then analyzes the video data again and makes new edits. In this case, too, the re-editing is done using video editing software, and the re-edited video is provided to the user.
[0115] Specific examples
[0116] 1. Video recording of a wedding reception
[0117] Cameras with wireless communication capabilities will be placed around the venue, and initial settings will be performed using a dedicated application.
[0118] The camera captures the bride and groom's entrance, guest speeches, toasts, and cake cutting moments in real time as H.264 format video streams and sends them to the server.
[0119] The server analyzes the collected data and uses an artificial intelligence model to perform facial recognition to identify the bride and groom and important guests. AI is also used to extract highlights based on emotional expressions and volume changes, and the footage is then edited in real time using video editing software.
[0120] Users (attendees) upload the video they have taken with their smartphones to the device using a dedicated upload application, and then transfer it to the server.
[0121] The server analyzes the uploaded footage and combines it with existing edited footage to create a richer final video.
[0122] The server then combines all the clips, adds transitions and titles, and generates the final edited video, which is then made available to the user via cloud storage.
[0123] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0124] Step 1:
[0125] The terminal installs wireless communication-enabled cameras at various locations in the venue. First, a dedicated setting application is launched and the camera's IP address, field of view, and installation location are entered. This information is then sent to the server and registered in a database.
[0126] Input: Camera IP address, viewing angle, installation location
[0127] Output: Camera information stored in the server database
[0128] Step 2:
[0129] The camera collects video and audio from the venue in real time as an H.264 video stream, and transmits this data to a server via wireless communication. The camera automatically establishes a connection and provides a live feed to a specified URL.
[0130] Input: Video and audio collected by a camera
[0131] Output: H.264 video stream (sent to server)
[0132] Step 3:
[0133] The server receives and analyzes the streaming data in real time. It then activates a facial recognition system to identify guests based on pre-registered group photos and video clips. Specifically, the server invokes an artificial intelligence model to perform facial recognition.
[0134] Input: H.264 format video stream
[0135] Output: Identified guest data
[0136] Step 4:
[0137] The server uses an AI model to analyze video and audio data and identify highlights based on emotional expressions and changes in volume. For example, it can extract important scenes by detecting changes in the volume of smiles or applause.
[0138] Input: Video stream after facial recognition
[0139] Output: Identified highlight scene data
[0140] Step 5:
[0141] The server performs real-time editing based on the identified highlight scenes. Using Adobe Premiere Pro API and FFmpeg, important scenes are extracted as clips and placed on the timeline. The edited footage is temporarily saved in an appropriate format.
[0142] Input: Identified highlight scene data
[0143] Output: Edited video clip (temporarily saved)
[0144] Step 6:
[0145] Users (attendees) upload video data taken with their own devices to the terminal using a dedicated upload application, and the terminal then sends this data to the server.
[0146] Input: Video data taken by the user
[0147] Output: Video data uploaded to the server
[0148] Step 7:
[0149] The server analyzes the received video data and integrates it with the existing highlight footage, inserting additional footage into the existing data at the appropriate time to create a more immersive final video.
[0150] Input: Uploaded user video data
[0151] Output: Final merged video data
[0152] Step 8:
[0153] The server then combines all the video data to create the final video. Specifically, it combines the edited clips, adds transitions and titles, and creates a timeline. The final video is saved in a standard format (MP4 or MOV) and uploaded to cloud storage.
[0154] Input: Final merged video data
[0155] Output: Final video file saved in cloud storage
[0156] Step 9:
[0157] The user sends a request to re-edit a specific scene to the server via their device. They enter details into a request form and send it to the server. The server then analyzes the video data again and performs a new edit. This is done using video editing software.
[0158] Input: Request for re-edit
[0159] Output: Re-edited video
[0160] Step 10:
[0161] The server saves the re-edited video back to cloud storage and provides the user with a shared link, allowing them to easily access the newly edited video.
[0162] Input: Re-edited video
[0163] Output: Video Recut files saved in cloud storage and a share link
[0164] (Application example 1)
[0165] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0166] In traditional store operations, analyzing customer behavior, providing efficient customer service, implementing security measures, and implementing effective marketing often consumes a significant amount of time and cost. It can also be difficult to grasp the situation in real time and take prompt action. This creates challenges that prevent efficient store operations and improved customer experience.
[0167] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0168] In this invention, the server includes a means for analyzing video and audio data obtained by the camera in real time to identify people's movement paths and areas where they are staying, a means for notifying detected information in real time and encouraging proactive responses, and a means for generating a heat map of customer traffic and displaying it on an interface accessible by a manager, thereby making it possible to improve the efficiency of store operations and customer satisfaction.
[0169] A "camera" is a device used to collect video and audio and transmit it to a server via wireless communication.
[0170] A "server" is a computer system that analyzes collected video and audio data and performs various processes such as identification, editing, and notification.
[0171] "Machine learning" is a technology that analyzes video and audio data to automatically recognize and identify specific patterns and highlight scenes.
[0172] "Facial recognition" is a technology for recognizing a person's face from collected video data and identifying that person.
[0173] A "highlight scene" refers to a particularly important or interesting moment or action in video or audio data.
[0174] "Wireless communication" is a technology that uses wireless LAN to send and receive data.
[0175] "User" refers to an individual or organization that shoots video data and uploads the data to a server using a terminal.
[0176] A "terminal" is a device (such as a smartphone or tablet) that a user uses to upload video data to a server.
[0177] "Traffic lines" refer to the routes that show how people move around the store.
[0178] A "stay area" refers to an area where a person or customer stays in a particular location for a certain period of time.
[0179] "Real-time notifications" are notifications sent instantly to prompt necessary action by immediately processing collected information.
[0180] "Proactive response" means predicting the situation and providing appropriate measures and services immediately.
[0181] A "heat map" is a visual display of the popularity and frequency of use of a particular area, using color coding based on collected data.
[0182] This invention is a system that uses multiple cameras to collect video and audio in real time, and a server analyzes and edits the data. This system can be applied to store operations, customer service, security measures, and marketing optimization.
[0183] Hardware and software used
[0184] Hardware:
[0185] Wireless LAN camera: Collects video and audio in real time (e.g., a typical network camera).
[0186] Server: Analyzes video and audio data and performs various processes such as identification, editing, and notification (e.g., cloud servers such as AWS EC2).
[0187] Smartphone: A device used by users to shoot and upload video data (e.g., iPhone, Android device).
[0188] Robots: Serve customers and guide them in stores (e.g., Pepper robot).
[0189] software:
[0190] Video analytics AI: Analyzes video data to identify customer movement patterns and areas of congestion (e.g., OpenCV, TensorFlow).
[0191] Speech recognition AI: Converts voice data into text and recognizes questions or requests (e.g., Google Speech-to-Text API).
[0192] Database: Stores and manages collected data (e.g., AWS RDS).
[0193] Real-time communication: Providing immediate notifications (e.g., WebSockets).
[0194] Streaming services: Real-time delivery of video data (e.g., AWS Kinesis).
[0195] System Overview
[0196] 1. Camera Initial Setup:
[0197] The camera is connected to the wireless LAN, and the terminal registers the camera's IP address, field of view, and installation location on the server.
[0198] 2. Video Collection and Transmission:
[0199] The camera collects video and audio from the entire venue in real time and transmits it to a server via wireless communication.
[0200] 3. Video and audio analysis:
[0201] The server receives the streaming data in real time and uses machine learning to analyze the video and audio data, identifying people through facial recognition and pinpointing their movement patterns and areas of concentration.
[0202] 4. Real-time notifications and actions:
[0203] The detected information is sent in real time to the store clerk's smartphone or the robot, encouraging proactive response.
[0204] 5. Heatmap and report generation:
[0205] Customer traffic heat maps are generated and displayed in the admin interface, and regular reports are generated based on the collected and analyzed data.
[0206] 6. Re-editing and final video creation:
[0207] The server re-edits specific scenes as needed and generates and serves the final video in response to the user's request.
[0208] Specific examples
[0209] Customer behavior analysis:
[0210] Cameras track customers' movements within the store and measure the time they spend in front of specific product areas, and the server then reports back to them the most popular products in the areas where customers spend the most time.
[0211] Prompt Sentence Examples
[0212] Customer behavior analysis:
[0213] - "Please identify customer movement patterns and areas where customers stay using this camera video data."
[0214] - "Please recognize the customer's question from this voice data."
[0215] This system will improve the efficiency of store operations and customer satisfaction, allowing for real-time situation assessment and rapid response.
[0216] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0217] Step 1:
[0218] Camera Defaults:
[0219] The terminal installs wireless LAN-enabled cameras in various areas of the store. The terminal registers the camera's IP address, field of view, and installation location on the server. The terminal receives the camera's physical installation location and setting information as input, and stores this information in a database as output.
[0220] Step 2:
[0221] Video collection and transmission:
[0222] The cameras collect video and audio from the entire venue in real time. This data is sent to the server via wireless communication. The system receives video and audio data from the cameras as input and sends it to the server via wireless communication. The output is treated as streaming data.
[0223] Step 3:
[0224] Video and audio analysis:
[0225] The server receives the streaming data in real time and analyzes it using video analysis AI and voice recognition AI. This allows it to identify customer movement patterns, areas where customers stay, personal identification using facial recognition, and the details of their questions and requests. It receives streaming data as input and stores the identified information (movement patterns, areas where customers stay, personal identification results, and questions) as output in a database.
[0226] Step 4:
[0227] Real-time notifications and actions:
[0228] The server notifies the detected information in real time to the store clerk's smartphone or the in-store robot, allowing the store clerk or robot to proactively respond to customers. The server receives the analysis results as input and sends notification data as output via WebSocket.
[0229] Step 5:
[0230] Generate a heatmap:
[0231] The server generates a heat map of customer traffic within the store based on data on customer movement and lingering areas. The generated heat map is displayed on an interface that can be accessed by administrators. The server receives data on movement and lingering areas as input, generates a heat map as output, and displays it on the interface.
[0232] Step 6:
[0233] Report Generation:
[0234] The server periodically generates reports based on the collected and analyzed data. The reports include customer movement patterns, areas where customers stay, questions, etc. It receives various data stored in the database as input, generates reports as output, and sends them to the administrator.
[0235] Step 7:
[0236] Re-editing and producing the final video:
[0237] The server re-edits specific scenes as needed and generates and provides the final video in response to the user's request. It receives the user's request and video data as input, generates the re-edited video as output, and provides it to the user.
[0238] In this way, real-time data collection and analysis, immediate notification and response, and visualization and reporting of analysis results will lead to more efficient store operations and improved customer satisfaction.
[0239] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0240] This invention combines a system that uses multiple cameras to collect and analyze video and audio in real time with an emotion engine. Below, we will explain the specific program processing of this system in natural language.
[0241] Camera Initial Settings
[0242] The terminal places wireless LAN-compatible cameras in various locations in the venue. The cameras are installed on the ceiling, on tabletops, and around tables, and registers information such as the camera's IP address, field of view, and installation location on the server.
[0243] Video collection and transmission
[0244] The camera begins collecting video and audio from the entire venue in real time and transmits it to a server via wireless LAN.
[0245] Video and audio analysis
[0246] The server receives the streaming data sent from the camera in real time, analyzes the video and audio, and primarily performs facial recognition to identify guests based on group photos and data acquired in advance.
[0247] Incorporating an emotion engine
[0248] The server uses an emotion engine to analyze the user's emotional state in real time based on the collected video and audio data. Based on this emotional state, highlight scenes are identified according to emotional expressions and volume changes in the video. For example, moments of cheering or crying are automatically extracted.
[0249] Video editing
[0250] The server edits the video in real time based on the identified highlight scenes, cutting out unnecessary parts and optimizing scene transitions. It also appropriately edits scenes that should be emphasized according to the emotional state identified by the emotion engine.
[0251] Attendee Video Integration
[0252] Users (attendees) upload video data taken with their smartphones or cameras to their devices. The devices then send the uploaded data to the server, which receives and analyzes it. The uploaded video data is integrated into existing highlight footage based on the analysis results of the emotion engine, creating a more moving and realistic video.
[0253] Generating the final video
[0254] The server generates the final video from all the video data, and the resulting video file is provided to the user in an easily accessible format.
[0255] Re-editing function
[0256] The user sends a request to the server via their device to re-edit a specific scene. For example, a specific request such as "I want more emphasis on the bride and groom's reaction" can be accepted. The server re-analyzes the video data based on the re-editing request and performs a new edit. The re-edited video is then provided to the user.
[0257] Specific examples
[0258] 1. Video recording of a wedding reception
[0259] The terminal places wireless LAN-compatible cameras at various locations throughout the venue and performs initial setup.
[0260] The camera captures the moment when the bride and groom appear, the guests' speeches, the toast, the cake cutting, and the expressions of the many guests in real time, and sends the video and audio data to the server.
[0261] The server analyzes the transmitted data and identifies the bride and groom and important guests using facial recognition, analyzes their emotional state using an emotion engine, and identifies highlight scenes based on emotional expressions and volume changes.
[0262] The server edits the video in real time based on the identified highlights. For example, the emotion engine detects the bride's tears of joy and emphasizes these scenes.
[0263] Users (attendees) upload videos they have taken with their smartphones, which are then received and analyzed by the server. The uploaded video is then integrated with existing edited footage based on the analysis results of the emotion engine, allowing for a richer reproduction of the bride and groom's reactions and the guests' emotional moments.
[0264] The server generates the final edited video and makes it ready to be served.
[0265] In this way, a system that combines an emotion engine enables high-quality video recording and editing that reflects the user's emotional state.
[0266] The processing flow will be explained below.
[0267] Step 1:
[0268] The terminal installs wireless LAN-compatible cameras in advance at various locations in the venue and registers information such as the camera's IP address, field of view, and installation location on the server.
[0269] Step 2:
[0270] The cameras begin collecting video and audio from the entire venue in real time and streaming it to a server via Wi-Fi.
[0271] Step 3:
[0272] The server receives the streaming data sent from the camera in real time and analyzes the video and audio.
[0273] Step 4:
[0274] The server performs facial recognition based on group photos and data acquired in advance to identify guests.
[0275] Step 5:
[0276] The server uses AI to analyze video and audio data, detecting emotional expressions and changes in volume to identify highlight scenes.
[0277] Step 6:
[0278] The server uses an emotion engine to analyze the user's emotional state in real time, and identifies important scenes in the video based on the emotional state.
[0279] Step 7:
[0280] The server then prioritizes editing the identified highlights and optimizes scene transitions by cutting out unnecessary parts, for example, by emphasizing touching moments between the bride and groom or scenes that elicit cheers of joy.
[0281] Step 8:
[0282] Users (attendees) upload video data taken with their smartphones or cameras to the device.
[0283] Step 9:
[0284] The device sends the uploaded video data of attendees to the server, which receives and analyzes it.
[0285] Step 10:
[0286] The server then integrates the uploaded attendee video data with existing highlight footage based on the analysis results of the emotion engine, creating a more moving and realistic video.
[0287] Step 11:
[0288] The server generates the final video from all the video data and provides the video file in an easily accessible format for the user.
[0289] Step 12:
[0290] The user sends a request to the server via their device to re-edit a specific scene, such as "I want more emphasis on the reaction of the bride and groom."
[0291] Step 13:
[0292] The server re-analyzes the video data based on the re-editing request and performs a new edit. The re-edited video is then provided to the user.
[0293] Example 2
[0294] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0295] Modern events require efficient recording and editing of many important moments. However, doing so manually is a time-consuming and labor-intensive task. Furthermore, integrating footage shot separately by attendees is tedious and does not necessarily produce natural-looking results. Furthermore, capturing emotional expressions and volume changes in the footage and identifying moving moments and important scenes is a difficult challenge. Therefore, there is a need for a system that can automatically collect, analyze, and edit footage in real time, and efficiently integrate additional footage from users.
[0296] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0297] In this invention, the server includes means for collecting video and audio from a camera in real time, means for transmitting the collected video and audio to the server via wireless communication, means for the server to perform individual identification using facial recognition technology, means for analyzing the video and audio using machine learning to identify important scenes, means for editing in real time based on the important scenes identified by the server, means for uploading video data shot by a user to the server using a terminal, means for integrating the uploaded user's video data with existing important scene videos, means for generating a final video using all the video data, means for re-editing based on a request to re-edit a specific scene, and means for outputting the generated final video and re-edited video. This makes it possible to automatically and efficiently collect, analyze, and edit video, and naturally integrate additional user video.
[0298] A "camera" is a device for collecting video and audio and capable of transmitting data to a server via a wireless communication network.
[0299] "Server" is a central processing unit for receiving, analyzing, and editing collected video and audio data.
[0300] "Facial recognition technology" refers to algorithms and techniques for identifying people from video data.
[0301] "Machine learning" is a technology that analyzes large amounts of data to learn patterns and characteristics and make predictions and distinctions.
[0302] An "important scene" refers to a particularly noteworthy or moving moment in a video, and is identified by emotional expressions or changes in volume.
[0303] "Real-time editing" is the process of editing scenes that are deemed important at the moment as video data is collected.
[0304] "User-recorded video data" refers to video files that event attendees record themselves using their smartphones or cameras and later upload to the system.
[0305] "Uploading" refers to the act of sending video data taken by a user to a server.
[0306] "Merge" is the process of combining uploaded footage with existing footage and key scenes.
[0307] The "final video" is a video file in the final format after all video data has been edited and compiled.
[0308] "Re-editing" is a process in which already edited video data is re-edited based on a specific user request.
[0309] "Generated final video or re-edited video" includes final version video files automatically generated by the system and re-edited video files based on user requests.
[0310] A "wireless communication enabled camera" is a camera that has the ability to transmit data using a wireless communication network.
[0311] A "communications network" is an infrastructure for transmitting and receiving data between multiple devices using wireless LAN or other wireless communication protocols.
[0312] "Emotional expression" refers to the psychological state and emotional expression obtained by analyzing facial expressions in the video.
[0313] "Volume change" refers to an increase or decrease in volume detected by analyzing the audio data in the video.
[0314] The present invention combines a system that utilizes multiple wireless communication enabled cameras to collect and analyze video and audio in real time with an emotion engine. Specific embodiments for implementing this system are described below.
[0315] First, wireless communication-enabled cameras are installed in the venue. The device (e.g., laptop or tablet PC) sets the camera's IP address, field of view, and installation location, and registers this data on the server. The cameras are installed on the ceiling, on tabletops, around tables, etc., and are positioned so that they cover the entire venue. At this point, the cameras and server are connected using a wireless communication network (such as wireless LAN).
[0316] Next, the camera starts collecting video and audio in real time. The camera also collects audio data using its built-in microphone and transmits this data wirelessly to the server. The server uses a high-performance CPU and GPU to receive and analyze the streaming data sent from the camera in real time.
[0317] The server first uses facial recognition technology (e.g., OpenCV or dlib) to identify individuals from the collected video data, then uses an algorithm to analyze the audio data, detecting specific keywords and volume changes. The identified data is then stored in a database.
[0318] Furthermore, the server uses an emotion engine (e.g., Emotion API or Affectiva) to analyze the user's emotional state from video and audio data. The emotion engine analyzes facial expressions and audio tones for each frame to identify emotional states such as smiling or crying faces. This allows highlight scenes (important scenes) to be automatically identified.
[0319] Based on the identified highlight scenes, the server edits the video in real time. For editing, editing software (e.g., FFmpeg) is used to cut out unnecessary parts and ensure smooth scene transitions. Scenes that should be emphasized are appropriately edited according to emotional expressions and volume changes.
[0320] In addition, users (attendees) can upload video data taken with their smartphones or cameras to their devices, which then send the uploaded data to the server. The server receives the data, analyzes it using an emotion engine, and integrates it into existing highlight footage. This creates a more moving and realistic video.
[0321] Finally, the server generates the final video from all the video data, and the completed video file is uploaded to cloud storage and provided to users in a format that is easy to access (e.g., MP4, AVI).
[0322] Furthermore, users can send requests to the server via their devices to re-edit specific scenes. For example, a specific request could be made to "emphasize the bride and groom's reactions more." The server then re-analyzes the original editing data based on the re-editing request and performs a new edit. The re-edited video is also uploaded to cloud storage and provided to the user.
[0323] Specific examples
[0324] 1. Video recording of a wedding reception
[0325] The terminals are used to place wireless communication-enabled cameras at various locations in the venue, and the initial settings are performed. The camera's IP address, field of view, and installation location are registered on the server.
[0326] The camera captures the moment when the bride and groom appear, the guests' speeches, the toast, the cake cutting, and the expressions of the many guests in real time, and sends the video and audio data to the server.
[0327] The server analyzes the transmitted data and uses facial recognition to identify the bride and groom and important guests, and an emotion engine to analyze their emotional state and identify highlight scenes based on emotional expressions and volume changes.
[0328] The server edits the video in real time based on the identified highlights. For example, the emotion engine detects the bride's tears of joy and emphasizes these scenes.
[0329] Users (attendees) upload videos they have taken with their smartphones, which are then received and analyzed by the server. The uploaded video is then integrated with existing edited footage based on the analysis results of the emotion engine, allowing for a richer reproduction of the bride and groom's reactions and the guests' emotional moments.
[0330] The server generates the final edited video and makes it ready to be served. The video file is uploaded to cloud storage in MP4 format and a download link is provided to the user.
[0331] Example prompts to input to the generative AI model
[0332] "I want to create a video that selects and edits the most touching scenes from a wedding reception."
[0333] "Create a highlight video that highlights the bride and groom's reactions and the guests' joyful moments."
[0334] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0335] Step 1: Initial camera setup
[0336] The terminal places wireless communication-enabled cameras in the venue. Specifically, it sets the camera's IP address, field of view, and installation location information, and registers this data on the server. For example, the terminal checks the network connection of each camera from a settings screen and enters the setting information. The input is the camera's physical installation location, field of view, and IP address. The output is the setting information stored in the server's database.
[0337] Step 2: Collect and transmit video and audio
[0338] The camera starts collecting video and audio in real time and sends them to the server via wireless communication. Specifically, the camera captures frames every second and collects audio data. The input is the video and audio acquired from the camera's sensor. The output is data sent to the server in streaming format.
[0339] Step 3: Video and audio analysis
[0340] The server receives and analyzes streaming data sent from the camera in real time. Individual identification is performed using facial recognition technology (e.g., OpenCV or dlib). The input is video data sent from the camera. Specifically, the captured frames are run through a facial recognition algorithm to detect the position of the face and its features. The output is information about the identified person and feature data. Audio data is also analyzed to detect specific keywords and volume changes. The input is audio data sent from the camera. The output is keyword detection information and volume change information as the analysis results.
[0341] Step 4: Incorporating the Emotion Engine
[0342] The server uses an emotion engine (e.g., Emotion API or Affectiva) to analyze the user's emotional state from video and audio data. The input is the video and audio feature data analyzed in the previous step. Specifically, the emotion engine analyzes facial expressions and audio tone for each frame to identify emotional states such as laughter or tears. The output is data related to the emotional state.
[0343] Step 5: Edit your footage
[0344] The server edits the video in real time based on the highlight scenes identified by the emotion engine. Specifically, it uses editing software (e.g., FFmpeg) to cut out unnecessary parts and smooth scene transitions. The input is the identified highlight scene information and the original video data. The output is the edited video data.
[0345] Step 6: Integrate attendee videos
[0346] Users (attendees) upload video data taken with their smartphones or cameras to their terminals. The terminals then send the uploaded data to the server. Specifically, the terminals manage the process of selecting video files and transferring them to the server. The input is the video data uploaded from the user's device. The output is additional video data stored on the server. The server analyzes the data using an emotion engine and integrates it into the existing highlight video. The input is the newly uploaded video data and the existing highlight video data. The output is the integrated video data.
[0347] Step 7: Generate the final video
[0348] The server generates the final video from all the video data. Specifically, editing software is used to combine all the scenes and create the final video file. The input is the merged video data. The output is the final video file (e.g., MP4 format). The generated file is uploaded to cloud storage so that users can access it.
[0349] Step 8: Re-edit function
[0350] The user sends a request to the server via their device to re-edit a specific scene. For example, a request may be made to "emphasize the bride and groom's reaction more." The server re-analyzes the original video data based on the re-editing request and performs a new edit. Specifically, the server re-processes the original video data according to the re-editing instructions and edits it to emphasize the necessary parts. The input is the re-editing request and the original video data. The output is a re-edited video file, which is also uploaded to cloud storage and a download link is provided to the user.
[0351] (Application example 2)
[0352] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0353] There is a need for a system that can analyze customer emotions in real time and provide appropriate feedback to store staff to provide more effective customer service and improve customer satisfaction. However, existing systems lack the means to analyze customer emotions and behavior in real time, making it difficult for staff to respond immediately.
[0354] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting video and audio from a camera in real time, means for transmitting the collected video and audio to the server via wireless communication, means for the server to identify people using biometric authentication, means for the server to analyze the video and audio using a generative AI model and identify key scenes, means for editing in real time based on the key scenes identified by the server, means for uploading video data taken by participants to the server using a terminal, means for integrating the uploaded video data of participants with existing key scene videos, means for generating a final video using all video data, means for re-editing based on a request to re-edit a specific scene, means for outputting the generated final video and re-edited video, and means for the smart device to analyze customer emotions in real time and provide feedback to staff. This makes it possible to analyze customer emotions in real time within the store and provide prompt and appropriate support.
[0355] A "camera" is a device that collects video and audio and converts it into electrical signals.
[0356] "Wireless communication" is a means of sending and receiving data without using cables, and is a technology that uses radio signals.
[0357] A "server" is a computer system that processes and stores data on a network and provides various services.
[0358] "Biometric authentication" is a technology that identifies individuals using biological characteristics such as their face, fingerprints, and irises.
[0359] A "generative AI model" is an algorithm or program that uses artificial intelligence to analyze data and generate new information or results.
[0360] "Video" is visual data generated by optical means and collected by a camera or other imaging device.
[0361] "Sound" refers to a sound signal transmitted by vibrations in the air, and is collected by a sound collection device such as a microphone.
[0362] An "important scene" refers to a scene or moment in video or audio data that is particularly noteworthy.
[0363] "Real-time" refers to data processing and analysis without delay, providing results almost instantly.
[0364] A "terminal" is a device that connects to a network and transmits, receives, and processes data.
[0365] "Integration" refers to combining multiple data or elements into one.
[0366] A "participant" is an individual or group who is actively involved in an event, activity, or project.
[0367] A "smart device" is an advanced electronic device that has communication capabilities and processes and presents information.
[0368] This invention describes a system that analyzes customer sentiment in brick-and-mortar stores and provides real-time feedback to store staff. The system is composed of a combination of cameras, audio collection devices, wireless communication connections, a server, smart devices, and a generative AI model.
[0369] Hardware and Software Configuration
[0370] Camera and audio collection device
[0371] Cameras are installed in various locations throughout the store and collect video and audio recordings of customers in real time. They have wireless communication capabilities and transmit video and audio data to a server. Specifically, store staff wear smart glasses (e.g., Google Glass or Vuzix M400) to collect customer interactions.
[0372] server
[0373] The server receives the collected video and audio data via wireless communication and analyzes it. This analysis involves identifying customers using biometric authentication technology and performing sentiment analysis using generative AI models (e.g., Affectiva, Microsoft Azure Cognitive Services). Furthermore, key scenes are identified based on the video and audio data, and editing is performed in real time.
[0374] Smart Devices
[0375] The smart device (smart glasses) displays the analysis results from the server, allowing store staff to instantly understand the customer's emotional state and respond appropriately.
[0376] Data processing and calculation
[0377] The server does the following:
[0378] 1. Receive video and audio data and identify customers using biometric technology.
[0379] 2. Use generative AI models to analyze customer sentiment in real time from video and audio data.
[0380] 3. Identify key scenes based on the customer's emotional state, edit in real time, and provide feedback to smart devices.
[0381] Specific examples
[0382] In practice, the following scenarios can be considered when using this system:
[0383] Situation 1: Dealing with customers in the fitting room
[0384] Example prompt sentence:
[0385] "If a staff member wearing glasses carefully observes a customer entering a fitting room and detects that the customer has a troubled expression."
[0386] What actually happens:
[0387] Smart glasses observe customers in fitting rooms.
[0388] The emotion engine analyzes the customer's facial expression to determine whether they are confused.
[0389] "The customer is in trouble" is displayed in real time on the smart glasses display.
[0390] Staff immediately went to the fitting room to provide the customer with appropriate assistance.
[0391] Situation 2: Serving customers at the counter
[0392] Example prompt sentence:
[0393] "When a staff member wearing glasses observes a customer who comes to a product counter, they detect that the customer has a satisfied expression."
[0394] What actually happens:
[0395] Smart glasses observe customers at the counter.
[0396] The emotion engine analyzes the customer's facial expression to determine whether they are satisfied.
[0397] "Customer is satisfied" is displayed in real time on the smart glasses display.
[0398] Staff will further enhance customer service, suggesting additional products and offering special services to customers.
[0399] This makes it possible to analyze customer sentiment in real time and provide appropriate support.
[0400] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0401] Step 1:
[0402] The cameras and audio collection devices begin to collect video and audio data in real time. This data captures how customers move around the store and what they say. The input is the video and audio captured by the cameras and microphones, and the output is the data transmitted wirelessly to the server.
[0403] Step 2:
[0404] The server receives and analyzes the data sent from the camera and audio collection device. Specifically, it uses a biometric authentication algorithm to identify customers from the video data. The input is the video and audio data sent from the camera and audio collection device, and the output is the ID information of the identified customer.
[0405] Step 3:
[0406] The server uses a generative AI model to analyze the customer's emotions from the received video and audio data, including facial expressions and tone of voice. The input is the identified customer's video and audio data, and the output is the customer's emotional state (e.g., interest, confusion, satisfaction, etc.).
[0407] Step 4:
[0408] The server identifies key moments in real time based on the customer's emotional state. Here, a generative AI model detects emotional peaks and changes and extracts specific events (e.g., excitement, confusion). The input is the customer's emotional state data, and the output is the timestamp and context of the identified key moments.
[0409] Step 5:
[0410] The server edits the video in real time based on the identified important scenes, cutting out unnecessary parts and emphasizing the important scenes. The input is the timestamp and context of the identified important scenes, and the output is the edited video clip.
[0411] Step 6:
[0412] The server sends the edited video clip to the smart device and provides feedback to the store staff in real time. The input is the edited video clip and the emotion analysis results, and the output is the feedback information displayed on the smart device's display.
[0413] Step 7:
[0414] Store staff wearing smart devices respond based on the feedback provided by the server. Specifically, if a customer's confusion is detected, the staff responds immediately and provides appropriate support. The input is the feedback information displayed on the smart device, and the output is the staff's immediate response.
[0415] Step 8:
[0416] The server stores all collected data and later analyzes it for long-term patterns and trends of customer behavior. The input is the entire video and audio data collected daily, and the output is regular reports and analysis results, providing information to support strategic decision-making.
[0417] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0418] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0419] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0420] [Second embodiment]
[0421] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0422] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0423] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0424] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0425] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0426] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0427] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0428] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0429] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0430] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0431] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0432] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0433] This invention is a system that uses multiple cameras to collect video and audio in real time, and a server analyzes and edits the video and audio. The specific program processing is explained below in natural language.
[0434] Camera Initial Settings
[0435] The terminal installs wireless LAN-compatible cameras in various locations in the venue. The cameras are placed on the ceiling, on tabletops, and around tables, and the terminal registers each camera's IP address, field of view, and installation location on the server.
[0436] Video collection and transmission
[0437] Cameras collect video and audio data from the entire venue in real time, and this data is sent to a server via wireless LAN.
[0438] Video and audio analysis
[0439] The server receives and analyzes the streaming data in real time. First, it uses facial recognition to identify guests based on group photos and video clips. Next, it uses AI to analyze the video and audio data and identify highlights based on emotional expressions and volume changes.
[0440] Video editing
[0441] The server then edits the identified highlights in real time, prioritizing important scenes such as the bride and groom's entrance, speeches, toasts, and cake cutting. The edited footage is then temporarily saved.
[0442] Attendee Video Integration
[0443] Users (attendees) upload video data taken with their smartphones or cameras to the device. The device then sends the uploaded data to the server, which receives and analyzes it. The received video data is integrated with existing highlight footage to create a more realistic video.
[0444] Generating the final video
[0445] The server generates the final video from all the video data, resulting in a complete video at the end of the event. The generated video file is then provided to users in an easily accessible format.
[0446] Re-editing function
[0447] The user sends a request to the server via their device to re-edit a specific scene. For example, if the user has a specific request, such as "I want more emphasis on the birthday cake," the server re-analyzes the video data and makes a new edit. The re-edited video is then provided to the user.
[0448] Specific examples
[0449] 1. Video recording of a wedding reception
[0450] The terminal places wireless LAN-compatible cameras at various locations throughout the venue and performs initial setup.
[0451] The camera captures the moment when the bride and groom enter, the guests give speeches, the toast, and the cake cutting in real time, and sends the video and audio data to the server.
[0452] The server analyzes the collected data and uses facial recognition to identify the bride and groom and important guests. Using AI, it extracts highlights based on emotional expressions and volume changes and edits them in real time.
[0453] Users (attendees) upload videos taken with their smartphones to the device and transfer them to the server.
[0454] The server then combines the uploaded footage with existing edited footage to create a richer video.
[0455] The server generates the final edited video and makes it ready to be served.
[0456] In this way, advanced video recording and editing can be achieved using AI and networks, without relying on the skills of cameramen or editors.
[0457] The processing flow will be explained below.
[0458] Step 1:
[0459] The terminal installs wireless LAN-compatible cameras in advance at various locations in the venue and registers information such as the camera's IP address, field of view, and installation location on the server.
[0460] Step 2:
[0461] The cameras begin collecting video and audio from the entire venue in real time and stream this data to a server via wireless LAN.
[0462] Step 3:
[0463] The server receives the streaming data sent from the camera in real time and begins analyzing the video and audio.
[0464] Step 4:
[0465] The server performs facial recognition based on group photos and data acquired in advance to identify guests.
[0466] Step 5:
[0467] The server uses AI to analyze video and audio, detecting emotional expressions and changes in volume to identify highlight scenes.
[0468] Step 6:
[0469] The server then prioritizes editing the identified highlight scenes, cutting out unnecessary parts to optimize scene transitions, and temporarily stores the edited footage.
[0470] Step 7:
[0471] Users (attendees) upload video data taken with their smartphones or cameras to the device.
[0472] Step 8:
[0473] The device sends the uploaded video data of attendees to the server, which receives and analyzes it.
[0474] Step 9:
[0475] The server then combines the footage uploaded by attendees with existing highlight footage to create a more immersive video.
[0476] Step 10:
[0477] The server generates the final video from all the video data and provides the video file in an easily accessible format for the user.
[0478] Step 11:
[0479] The user sends a request to the server via their device to re-edit a specific scene.
[0480] Step 12:
[0481] The server re-analyzes the video based on the re-editing request, performs a new edit, and provides the re-edited video to the user.
[0482] Example 1
[0483] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0484] Conventional video recording systems rely on manual processes for collecting, analyzing, and editing video and audio, resulting in limitations in real-time performance and accuracy. This makes efficient, high-quality video editing particularly difficult in situations where large amounts of data must be collected in a short period of time, such as events and meetings. It is also cumbersome for users to easily integrate and re-edit footage they have shot themselves. There is a demand for a system that can solve these issues and achieve efficient, high-precision video recording and editing.
[0485] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0486] In this invention, the server includes means for collecting video and audio from a camera in real time, means for transmitting the collected video and audio to the server via wireless communication, means for the server to identify guests using facial recognition, means for the server to analyze the video and audio using artificial intelligence and identify highlight scenes, means for editing in real time based on the highlight scenes identified by the server, means for uploading video data taken by attendees to the server using a terminal, means for integrating the uploaded video data of attendees with existing highlight videos, means for generating a final video using all the video data, means for re-editing based on a request to re-edit specific scenes, and means for outputting the generated final video and re-edited video, thereby enabling efficient and high-quality video recording and editing.
[0487] A "camera" is a photographic device that collects video and audio from within the venue in real time and transmits them to a server.
[0488] "Wireless communication" is a general term for technology that uses radio waves to send and receive data.
[0489] A "server" is a computer system that receives data from multiple cameras and devices via a network and analyzes, edits, saves, and outputs the data.
[0490] "Facial recognition" is a technology that analyzes faces contained in video data and identifies individuals.
[0491] "Artificial intelligence" is a data analysis technology that uses machine learning and deep learning to analyze video based on emotional expressions, volume changes, etc.
[0492] "Highlight scenes" refer to important moments or scenes with strong emotional expressions within video data.
[0493] "Real-time editing" refers to the process of collecting video data and simultaneously analyzing and editing it to create a format that can be output immediately.
[0494] An "attendee" is a person attending an event, meeting, etc., including anyone who captures video or audio on their own device.
[0495] "Terminals" are devices used by attendees to upload video data to the server. These devices include smartphones and tablets.
[0496] "Integration" refers to the process of analyzing multiple video data sets and editing and combining them into a series of flows.
[0497] The "final video" is the completed video edited based on all the collected video data.
[0498] "Re-editing" refers to the process of adding new edits to already edited video data based on a request from a user.
[0499] "Generation" refers to the process of creating the final video from the collected data.
[0500] "Output" refers to providing the generated Final Video or Re-Edited Video in a user-accessible format.
[0501] The present invention is a system that uses multiple cameras to collect video and audio in real time, and a server analyzes and edits the collected video and audio. Specifically, the system is implemented according to the following steps:
[0502] First, wireless communication-enabled cameras are installed in various locations throughout the venue. The cameras are placed on the ceiling, on tabletops, and around tables, and the IP address, field of view, and installation location of each camera are registered on the server by the terminal. At this time, a dedicated setting application is used to save the camera information in the server database.
[0503] Next, the camera collects video and audio from the entire venue in real time. The collected data is sent to the server via wireless communication as an H.264 video stream. The camera automatically starts transmitting data and provides a live feed to a URL specified by the server.
[0504] The server receives and analyzes the streaming data in real time. First, it performs facial recognition based on group photos and pre-collected video clips to identify guests. Specifically, the server calls an AI model to perform facial recognition and identify individual guests. Next, it uses the AI model to analyze the video and audio data and identify highlight scenes based on emotional expressions and volume changes. For example, it prioritizes the extraction of scenes featuring the bride and groom or scenes with moving speeches.
[0505] The server then edits the identified highlights in real time. Specifically, the server uses video editing software to edit the video. Important scenes, such as the bride and groom's entrance, the toast, and the cake cutting, are extracted as clips and placed on the timeline. The edited video is then temporarily saved in an appropriate format.
[0506] Users (attendees) upload video data they have taken with their own devices to their terminals. This data is then sent to the server using a dedicated upload application. The video data received by the server is synchronized and integrated with existing highlight footage. The server analyzes this data and inserts additional footage at the appropriate times to create a more immersive final video.
[0507] The server then combines all the video data to create the final video. Specifically, it combines each edited clip and organizes the overall timeline. The video format is a common format such as MP4 or MOV. The final video file is then uploaded to a cloud storage service, and a shared link is provided for users to easily access.
[0508] It is also possible for users to send requests to the server via their device to re-edit specific scenes. For example, if a user has a specific request, such as "I want more emphasis on the toast scene," the user enters the details in a dedicated form and sends it to the server. The server then analyzes the video data again and makes new edits. In this case, too, the re-editing is done using video editing software, and the re-edited video is provided to the user.
[0509] Specific examples
[0510] 1. Video recording of a wedding reception
[0511] Cameras with wireless communication capabilities will be placed around the venue, and initial settings will be performed using a dedicated application.
[0512] The camera captures the bride and groom's entrance, guest speeches, toasts, and cake cutting moments in real time as H.264 format video streams and sends them to the server.
[0513] The server analyzes the collected data and uses an artificial intelligence model to perform facial recognition to identify the bride and groom and important guests. AI is also used to extract highlights based on emotional expressions and volume changes, and the footage is then edited in real time using video editing software.
[0514] Users (attendees) upload the video they have taken with their smartphones to the device using a dedicated upload application, and then transfer it to the server.
[0515] The server analyzes the uploaded footage and combines it with existing edited footage to create a richer final video.
[0516] The server then combines all the clips, adds transitions and titles, and generates the final edited video, which is then made available to the user via cloud storage.
[0517] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0518] Step 1:
[0519] The terminal installs wireless communication-enabled cameras at various locations in the venue. First, a dedicated setting application is launched and the camera's IP address, field of view, and installation location are entered. This information is then sent to the server and registered in a database.
[0520] Input: Camera IP address, viewing angle, installation location
[0521] Output: Camera information stored in the server database
[0522] Step 2:
[0523] The camera collects video and audio from the venue in real time as an H.264 video stream, and transmits this data to a server via wireless communication. The camera automatically establishes a connection and provides a live feed to a specified URL.
[0524] Input: Video and audio collected by a camera
[0525] Output: H.264 video stream (sent to server)
[0526] Step 3:
[0527] The server receives and analyzes the streaming data in real time. It then activates a facial recognition system to identify guests based on pre-registered group photos and video clips. Specifically, the server invokes an artificial intelligence model to perform facial recognition.
[0528] Input: H.264 format video stream
[0529] Output: Identified guest data
[0530] Step 4:
[0531] The server uses an AI model to analyze video and audio data and identify highlights based on emotional expressions and changes in volume. For example, it can extract important scenes by detecting changes in the volume of smiles or applause.
[0532] Input: Video stream after facial recognition
[0533] Output: Identified highlight scene data
[0534] Step 5:
[0535] The server performs real-time editing based on the identified highlight scenes. Using Adobe Premiere Pro API and FFmpeg, important scenes are extracted as clips and placed on the timeline. The edited footage is temporarily saved in an appropriate format.
[0536] Input: Identified highlight scene data
[0537] Output: Edited video clip (temporarily saved)
[0538] Step 6:
[0539] Users (attendees) upload video data taken with their own devices to the terminal using a dedicated upload application, and the terminal then sends this data to the server.
[0540] Input: Video data taken by the user
[0541] Output: Video data uploaded to the server
[0542] Step 7:
[0543] The server analyzes the received video data and integrates it with the existing highlight footage, inserting additional footage into the existing data at the appropriate time to create a more immersive final video.
[0544] Input: Uploaded user video data
[0545] Output: Final merged video data
[0546] Step 8:
[0547] The server then combines all the video data to create the final video. Specifically, it combines the edited clips, adds transitions and titles, and creates a timeline. The final video is saved in a standard format (MP4 or MOV) and uploaded to cloud storage.
[0548] Input: Final merged video data
[0549] Output: Final video file saved in cloud storage
[0550] Step 9:
[0551] The user sends a request to re-edit a specific scene to the server via their device. They enter details into a request form and send it to the server. The server then analyzes the video data again and performs a new edit. This is done using video editing software.
[0552] Input: Request for re-edit
[0553] Output: Re-edited video
[0554] Step 10:
[0555] The server saves the re-edited video back to cloud storage and provides the user with a shared link, allowing them to easily access the newly edited video.
[0556] Input: Re-edited video
[0557] Output: Video Recut files saved in cloud storage and a share link
[0558] (Application example 1)
[0559] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0560] In traditional store operations, analyzing customer behavior, providing efficient customer service, implementing security measures, and implementing effective marketing often consumes a significant amount of time and cost. It can also be difficult to grasp the situation in real time and take prompt action. This creates challenges that prevent efficient store operations and improved customer experience.
[0561] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0562] In this invention, the server includes a means for analyzing video and audio data obtained by the camera in real time to identify people's movement paths and areas where they are staying, a means for notifying detected information in real time and encouraging proactive responses, and a means for generating a heat map of customer traffic and displaying it on an interface accessible by a manager, thereby making it possible to improve the efficiency of store operations and customer satisfaction.
[0563] A "camera" is a device used to collect video and audio and transmit it to a server via wireless communication.
[0564] A "server" is a computer system that analyzes collected video and audio data and performs various processes such as identification, editing, and notification.
[0565] "Machine learning" is a technology that analyzes video and audio data to automatically recognize and identify specific patterns and highlight scenes.
[0566] "Facial recognition" is a technology for recognizing a person's face from collected video data and identifying that person.
[0567] A "highlight scene" refers to a particularly important or interesting moment or action in video or audio data.
[0568] "Wireless communication" is a technology that uses wireless LAN to send and receive data.
[0569] "User" refers to an individual or organization that shoots video data and uploads the data to a server using a terminal.
[0570] A "terminal" is a device (such as a smartphone or tablet) that a user uses to upload video data to a server.
[0571] "Traffic lines" refer to the routes that show how people move around the store.
[0572] A "stay area" refers to an area where a person or customer stays in a particular location for a certain period of time.
[0573] "Real-time notifications" are notifications sent instantly to prompt necessary action by immediately processing collected information.
[0574] "Proactive response" means predicting the situation and providing appropriate measures and services immediately.
[0575] A "heat map" is a visual display of the popularity and frequency of use of a particular area, using color coding based on collected data.
[0576] This invention is a system that uses multiple cameras to collect video and audio in real time, and a server analyzes and edits the data. This system can be applied to store operations, customer service, security measures, and marketing optimization.
[0577] Hardware and software used
[0578] Hardware:
[0579] Wireless LAN camera: Collects video and audio in real time (e.g., a typical network camera).
[0580] Server: Analyzes video and audio data and performs various processes such as identification, editing, and notification (e.g., cloud servers such as AWS EC2).
[0581] Smartphone: A device used by users to shoot and upload video data (e.g., iPhone, Android device).
[0582] Robots: Serve customers and guide them in stores (e.g., Pepper robot).
[0583] software:
[0584] Video analytics AI: Analyzes video data to identify customer movement patterns and areas of congestion (e.g., OpenCV, TensorFlow).
[0585] Speech recognition AI: Converts voice data into text and recognizes questions or requests (e.g., Google Speech-to-Text API).
[0586] Database: Stores and manages collected data (e.g., AWS RDS).
[0587] Real-time communication: Providing immediate notifications (e.g., WebSockets).
[0588] Streaming services: Real-time delivery of video data (e.g., AWS Kinesis).
[0589] System Overview
[0590] 1. Camera Initial Setup:
[0591] The camera is connected to the wireless LAN, and the terminal registers the camera's IP address, field of view, and installation location on the server.
[0592] 2. Video Collection and Transmission:
[0593] The camera collects video and audio from the entire venue in real time and transmits it to a server via wireless communication.
[0594] 3. Video and audio analysis:
[0595] The server receives the streaming data in real time and uses machine learning to analyze the video and audio data, identifying people through facial recognition and pinpointing their movement patterns and areas of concentration.
[0596] 4. Real-time notifications and actions:
[0597] The detected information is sent in real time to the store clerk's smartphone or the robot, encouraging proactive response.
[0598] 5. Heatmap and report generation:
[0599] Customer traffic heat maps are generated and displayed in the admin interface, and regular reports are generated based on the collected and analyzed data.
[0600] 6. Re-editing and final video creation:
[0601] The server re-edits specific scenes as needed and generates and serves the final video in response to the user's request.
[0602] Specific examples
[0603] Customer behavior analysis:
[0604] Cameras track customers' movements within the store and measure the time they spend in front of specific product areas, and the server then reports back to them the most popular products in the areas where customers spend the most time.
[0605] Prompt Sentence Examples
[0606] Customer behavior analysis:
[0607] - "Please identify customer movement patterns and areas where customers stay using this camera video data."
[0608] - "Please recognize the customer's question from this voice data."
[0609] This system will improve the efficiency of store operations and customer satisfaction, allowing for real-time situation assessment and rapid response.
[0610] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0611] Step 1:
[0612] Camera Defaults:
[0613] The terminal installs wireless LAN-enabled cameras in various areas of the store. The terminal registers the camera's IP address, field of view, and installation location on the server. The terminal receives the camera's physical installation location and setting information as input, and stores this information in a database as output.
[0614] Step 2:
[0615] Video collection and transmission:
[0616] The cameras collect video and audio from the entire venue in real time. This data is sent to the server via wireless communication. The system receives video and audio data from the cameras as input and sends it to the server via wireless communication. The output is treated as streaming data.
[0617] Step 3:
[0618] Video and audio analysis:
[0619] The server receives the streaming data in real time and analyzes it using video analysis AI and voice recognition AI. This allows it to identify customer movement patterns, areas where customers stay, personal identification using facial recognition, and the details of their questions and requests. It receives streaming data as input and stores the identified information (movement patterns, areas where customers stay, personal identification results, and questions) as output in a database.
[0620] Step 4:
[0621] Real-time notifications and actions:
[0622] The server notifies the detected information in real time to the store clerk's smartphone or the in-store robot, allowing the store clerk or robot to proactively respond to customers. The server receives the analysis results as input and sends notification data as output via WebSocket.
[0623] Step 5:
[0624] Generate a heatmap:
[0625] The server generates a heat map of customer traffic within the store based on data on customer movement and lingering areas. The generated heat map is displayed on an interface that can be accessed by administrators. The server receives data on movement and lingering areas as input, generates a heat map as output, and displays it on the interface.
[0626] Step 6:
[0627] Report Generation:
[0628] The server periodically generates reports based on the collected and analyzed data. The reports include customer movement patterns, areas where customers stay, questions, etc. It receives various data stored in the database as input, generates reports as output, and sends them to the administrator.
[0629] Step 7:
[0630] Re-editing and producing the final video:
[0631] The server re-edits specific scenes as needed and generates and provides the final video in response to the user's request. It receives the user's request and video data as input, generates the re-edited video as output, and provides it to the user.
[0632] In this way, real-time data collection and analysis, immediate notification and response, and visualization and reporting of analysis results will lead to more efficient store operations and improved customer satisfaction.
[0633] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0634] This invention combines a system that uses multiple cameras to collect and analyze video and audio in real time with an emotion engine. Below, we will explain the specific program processing of this system in natural language.
[0635] Camera Initial Settings
[0636] The terminal places wireless LAN-compatible cameras in various locations in the venue. The cameras are installed on the ceiling, on tabletops, and around tables, and registers information such as the camera's IP address, field of view, and installation location on the server.
[0637] Video collection and transmission
[0638] The camera begins collecting video and audio from the entire venue in real time and transmits it to a server via wireless LAN.
[0639] Video and audio analysis
[0640] The server receives the streaming data sent from the camera in real time, analyzes the video and audio, and primarily performs facial recognition to identify guests based on group photos and data acquired in advance.
[0641] Incorporating an emotion engine
[0642] The server uses an emotion engine to analyze the user's emotional state in real time based on the collected video and audio data. Based on this emotional state, highlight scenes are identified according to emotional expressions and volume changes in the video. For example, moments of cheering or crying are automatically extracted.
[0643] Video editing
[0644] The server edits the video in real time based on the identified highlight scenes, cutting out unnecessary parts and optimizing scene transitions. It also appropriately edits scenes that should be emphasized according to the emotional state identified by the emotion engine.
[0645] Attendee Video Integration
[0646] Users (attendees) upload video data taken with their smartphones or cameras to their devices. The devices then send the uploaded data to the server, which receives and analyzes it. The uploaded video data is integrated into existing highlight footage based on the analysis results of the emotion engine, creating a more moving and realistic video.
[0647] Generating the final video
[0648] The server generates the final video from all the video data, and the resulting video file is provided to the user in an easily accessible format.
[0649] Re-editing function
[0650] The user sends a request to the server via their device to re-edit a specific scene. For example, a specific request such as "I want more emphasis on the bride and groom's reaction" can be accepted. The server re-analyzes the video data based on the re-editing request and performs a new edit. The re-edited video is then provided to the user.
[0651] Specific examples
[0652] 1. Video recording of a wedding reception
[0653] The terminal places wireless LAN-compatible cameras at various locations throughout the venue and performs initial setup.
[0654] The camera captures the moment when the bride and groom appear, the guests' speeches, the toast, the cake cutting, and the expressions of the many guests in real time, and sends the video and audio data to the server.
[0655] The server analyzes the transmitted data and identifies the bride and groom and important guests using facial recognition, analyzes their emotional state using an emotion engine, and identifies highlight scenes based on emotional expressions and volume changes.
[0656] The server edits the video in real time based on the identified highlights. For example, the emotion engine detects the bride's tears of joy and emphasizes these scenes.
[0657] Users (attendees) upload videos they have taken with their smartphones, which are then received and analyzed by the server. The uploaded video is then integrated with existing edited footage based on the analysis results of the emotion engine, allowing for a richer reproduction of the bride and groom's reactions and the guests' emotional moments.
[0658] The server generates the final edited video and makes it ready to be served.
[0659] In this way, a system that combines an emotion engine enables high-quality video recording and editing that reflects the user's emotional state.
[0660] The processing flow will be explained below.
[0661] Step 1:
[0662] The terminal installs wireless LAN-compatible cameras in advance at various locations in the venue and registers information such as the camera's IP address, field of view, and installation location on the server.
[0663] Step 2:
[0664] The cameras begin collecting video and audio from the entire venue in real time and streaming it to a server via Wi-Fi.
[0665] Step 3:
[0666] The server receives the streaming data sent from the camera in real time and analyzes the video and audio.
[0667] Step 4:
[0668] The server performs facial recognition based on group photos and data acquired in advance to identify guests.
[0669] Step 5:
[0670] The server uses AI to analyze video and audio data, detecting emotional expressions and changes in volume to identify highlight scenes.
[0671] Step 6:
[0672] The server uses an emotion engine to analyze the user's emotional state in real time, and identifies important scenes in the video based on the emotional state.
[0673] Step 7:
[0674] The server then prioritizes editing the identified highlights and optimizes scene transitions by cutting out unnecessary parts, for example, by emphasizing touching moments between the bride and groom or scenes that elicit cheers of joy.
[0675] Step 8:
[0676] Users (attendees) upload video data taken with their smartphones or cameras to the device.
[0677] Step 9:
[0678] The device sends the uploaded video data of attendees to the server, which receives and analyzes it.
[0679] Step 10:
[0680] The server then integrates the uploaded attendee video data with existing highlight footage based on the analysis results of the emotion engine, creating a more moving and realistic video.
[0681] Step 11:
[0682] The server generates the final video from all the video data and provides the video file in an easily accessible format for the user.
[0683] Step 12:
[0684] The user sends a request to the server via their device to re-edit a specific scene, such as "I want more emphasis on the reaction of the bride and groom."
[0685] Step 13:
[0686] The server re-analyzes the video data based on the re-editing request and performs a new edit. The re-edited video is then provided to the user.
[0687] Example 2
[0688] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0689] Modern events require efficient recording and editing of many important moments. However, doing so manually is a time-consuming and labor-intensive task. Furthermore, integrating footage shot separately by attendees is tedious and does not necessarily produce natural-looking results. Furthermore, capturing emotional expressions and volume changes in the footage and identifying moving moments and important scenes is a difficult challenge. Therefore, there is a need for a system that can automatically collect, analyze, and edit footage in real time, and efficiently integrate additional footage from users.
[0690] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0691] In this invention, the server includes means for collecting video and audio from a camera in real time, means for transmitting the collected video and audio to the server via wireless communication, means for the server to perform individual identification using facial recognition technology, means for analyzing the video and audio using machine learning to identify important scenes, means for editing in real time based on the important scenes identified by the server, means for uploading video data shot by a user to the server using a terminal, means for integrating the uploaded user's video data with existing important scene videos, means for generating a final video using all the video data, means for re-editing based on a request to re-edit a specific scene, and means for outputting the generated final video and re-edited video. This makes it possible to automatically and efficiently collect, analyze, and edit video, and naturally integrate additional user video.
[0692] A "camera" is a device for collecting video and audio and capable of transmitting data to a server via a wireless communication network.
[0693] "Server" is a central processing unit for receiving, analyzing, and editing collected video and audio data.
[0694] "Facial recognition technology" refers to algorithms and techniques for identifying people from video data.
[0695] "Machine learning" is a technology that analyzes large amounts of data to learn patterns and characteristics and make predictions and distinctions.
[0696] An "important scene" refers to a particularly noteworthy or moving moment in a video, and is identified by emotional expressions or changes in volume.
[0697] "Real-time editing" is the process of editing scenes that are deemed important at the moment as video data is collected.
[0698] "User-recorded video data" refers to video files that event attendees record themselves using their smartphones or cameras and later upload to the system.
[0699] "Uploading" refers to the act of sending video data taken by a user to a server.
[0700] "Merge" is the process of combining uploaded footage with existing footage and key scenes.
[0701] The "final video" is a video file in the final format after all video data has been edited and compiled.
[0702] "Re-editing" is a process in which already edited video data is re-edited based on a specific user request.
[0703] "Generated final video or re-edited video" includes final version video files automatically generated by the system and re-edited video files based on user requests.
[0704] A "wireless communication enabled camera" is a camera that has the ability to transmit data using a wireless communication network.
[0705] A "communications network" is an infrastructure for transmitting and receiving data between multiple devices using wireless LAN or other wireless communication protocols.
[0706] "Emotional expression" refers to the psychological state and emotional expression obtained by analyzing facial expressions in the video.
[0707] "Volume change" refers to an increase or decrease in volume detected by analyzing the audio data in the video.
[0708] The present invention combines a system that utilizes multiple wireless communication enabled cameras to collect and analyze video and audio in real time with an emotion engine. Specific embodiments for implementing this system are described below.
[0709] First, wireless communication-enabled cameras are installed in the venue. The device (e.g., laptop or tablet PC) sets the camera's IP address, field of view, and installation location, and registers this data on the server. The cameras are installed on the ceiling, on tabletops, around tables, etc., and are positioned so that they cover the entire venue. At this point, the cameras and server are connected using a wireless communication network (such as wireless LAN).
[0710] Next, the camera starts collecting video and audio in real time. The camera also collects audio data using its built-in microphone and transmits this data wirelessly to the server. The server uses a high-performance CPU and GPU to receive and analyze the streaming data sent from the camera in real time.
[0711] The server first uses facial recognition technology (e.g., OpenCV or dlib) to identify individuals from the collected video data, then uses an algorithm to analyze the audio data, detecting specific keywords and volume changes. The identified data is then stored in a database.
[0712] Furthermore, the server uses an emotion engine (e.g., Emotion API or Affectiva) to analyze the user's emotional state from video and audio data. The emotion engine analyzes facial expressions and audio tones for each frame to identify emotional states such as smiling or crying faces. This allows highlight scenes (important scenes) to be automatically identified.
[0713] Based on the identified highlight scenes, the server edits the video in real time. For editing, editing software (e.g., FFmpeg) is used to cut out unnecessary parts and ensure smooth scene transitions. Scenes that should be emphasized are appropriately edited according to emotional expressions and volume changes.
[0714] In addition, users (attendees) can upload video data taken with their smartphones or cameras to their devices, which then send the uploaded data to the server. The server receives the data, analyzes it using an emotion engine, and integrates it into existing highlight footage. This creates a more moving and realistic video.
[0715] Finally, the server generates the final video from all the video data, and the completed video file is uploaded to cloud storage and provided to users in a format that is easy to access (e.g., MP4, AVI).
[0716] Furthermore, users can send requests to the server via their devices to re-edit specific scenes. For example, a specific request could be made to "emphasize the bride and groom's reactions more." The server then re-analyzes the original editing data based on the re-editing request and performs a new edit. The re-edited video is also uploaded to cloud storage and provided to the user.
[0717] Specific examples
[0718] 1. Video recording of a wedding reception
[0719] The terminals are used to place wireless communication-enabled cameras at various locations in the venue, and the initial settings are performed. The camera's IP address, field of view, and installation location are registered on the server.
[0720] The camera captures the moment when the bride and groom appear, the guests' speeches, the toast, the cake cutting, and the expressions of the many guests in real time, and sends the video and audio data to the server.
[0721] The server analyzes the transmitted data and uses facial recognition to identify the bride and groom and important guests, and an emotion engine to analyze their emotional state and identify highlight scenes based on emotional expressions and volume changes.
[0722] The server edits the video in real time based on the identified highlights. For example, the emotion engine detects the bride's tears of joy and emphasizes these scenes.
[0723] Users (attendees) upload videos they have taken with their smartphones, which are then received and analyzed by the server. The uploaded video is then integrated with existing edited footage based on the analysis results of the emotion engine, allowing for a richer reproduction of the bride and groom's reactions and the guests' emotional moments.
[0724] The server generates the final edited video and makes it ready to be served. The video file is uploaded to cloud storage in MP4 format and a download link is provided to the user.
[0725] Example prompts to input to the generative AI model
[0726] "I want to create a video that selects and edits the most touching scenes from a wedding reception."
[0727] "Create a highlight video that highlights the bride and groom's reactions and the guests' joyful moments."
[0728] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0729] Step 1: Initial camera setup
[0730] The terminal places wireless communication-enabled cameras in the venue. Specifically, it sets the camera's IP address, field of view, and installation location information, and registers this data on the server. For example, the terminal checks the network connection of each camera from a settings screen and enters the setting information. The input is the camera's physical installation location, field of view, and IP address. The output is the setting information stored in the server's database.
[0731] Step 2: Collect and transmit video and audio
[0732] The camera starts collecting video and audio in real time and sends them to the server via wireless communication. Specifically, the camera captures frames every second and collects audio data. The input is the video and audio acquired from the camera's sensor. The output is data sent to the server in streaming format.
[0733] Step 3: Video and audio analysis
[0734] The server receives and analyzes streaming data sent from the camera in real time. Individual identification is performed using facial recognition technology (e.g., OpenCV or dlib). The input is video data sent from the camera. Specifically, the captured frames are run through a facial recognition algorithm to detect the position of the face and its features. The output is information about the identified person and feature data. Audio data is also analyzed to detect specific keywords and volume changes. The input is audio data sent from the camera. The output is keyword detection information and volume change information as the analysis results.
[0735] Step 4: Incorporating the Emotion Engine
[0736] The server uses an emotion engine (e.g., Emotion API or Affectiva) to analyze the user's emotional state from video and audio data. The input is the video and audio feature data analyzed in the previous step. Specifically, the emotion engine analyzes facial expressions and audio tone for each frame to identify emotional states such as laughter or tears. The output is data related to the emotional state.
[0737] Step 5: Edit your footage
[0738] The server edits the video in real time based on the highlight scenes identified by the emotion engine. Specifically, it uses editing software (e.g., FFmpeg) to cut out unnecessary parts and smooth scene transitions. The input is the identified highlight scene information and the original video data. The output is the edited video data.
[0739] Step 6: Integrate attendee videos
[0740] Users (attendees) upload video data taken with their smartphones or cameras to their terminals. The terminals then send the uploaded data to the server. Specifically, the terminals manage the process of selecting video files and transferring them to the server. The input is the video data uploaded from the user's device. The output is additional video data stored on the server. The server analyzes the data using an emotion engine and integrates it into the existing highlight video. The input is the newly uploaded video data and the existing highlight video data. The output is the integrated video data.
[0741] Step 7: Generate the final video
[0742] The server generates the final video from all the video data. Specifically, editing software is used to combine all the scenes and create the final video file. The input is the merged video data. The output is the final video file (e.g., MP4 format). The generated file is uploaded to cloud storage so that users can access it.
[0743] Step 8: Re-edit function
[0744] The user sends a request to the server via their device to re-edit a specific scene. For example, a request may be made to "emphasize the bride and groom's reaction more." The server re-analyzes the original video data based on the re-editing request and performs a new edit. Specifically, the server re-processes the original video data according to the re-editing instructions and edits it to emphasize the necessary parts. The input is the re-editing request and the original video data. The output is a re-edited video file, which is also uploaded to cloud storage and a download link is provided to the user.
[0745] (Application example 2)
[0746] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0747] There is a need for a system that can analyze customer emotions in real time and provide appropriate feedback to store staff to provide more effective customer service and improve customer satisfaction. However, existing systems lack the means to analyze customer emotions and behavior in real time, making it difficult for staff to respond immediately.
[0748] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting video and audio from a camera in real time, means for transmitting the collected video and audio to the server via wireless communication, means for the server to identify people using biometric authentication, means for the server to analyze the video and audio using a generative AI model and identify key scenes, means for editing in real time based on the key scenes identified by the server, means for uploading video data taken by participants to the server using a terminal, means for integrating the uploaded video data of participants with existing key scene videos, means for generating a final video using all video data, means for re-editing based on a request to re-edit a specific scene, means for outputting the generated final video and re-edited video, and means for the smart device to analyze customer emotions in real time and provide feedback to staff. This makes it possible to analyze customer emotions in real time within the store and provide prompt and appropriate support.
[0749] A "camera" is a device that collects video and audio and converts it into electrical signals.
[0750] "Wireless communication" is a means of sending and receiving data without using cables, and is a technology that uses radio signals.
[0751] A "server" is a computer system that processes and stores data on a network and provides various services.
[0752] "Biometric authentication" is a technology that identifies individuals using biological characteristics such as their face, fingerprints, and irises.
[0753] A "generative AI model" is an algorithm or program that uses artificial intelligence to analyze data and generate new information or results.
[0754] "Video" is visual data generated by optical means and collected by a camera or other imaging device.
[0755] "Sound" refers to a sound signal transmitted by vibrations in the air, and is collected by a sound collection device such as a microphone.
[0756] An "important scene" refers to a scene or moment in video or audio data that is particularly noteworthy.
[0757] "Real-time" refers to data processing and analysis without delay, providing results almost instantly.
[0758] A "terminal" is a device that connects to a network and transmits, receives, and processes data.
[0759] "Integration" refers to combining multiple data or elements into one.
[0760] A "participant" is an individual or group who is actively involved in an event, activity, or project.
[0761] A "smart device" is an advanced electronic device that has communication capabilities and processes and presents information.
[0762] This invention describes a system that analyzes customer sentiment in brick-and-mortar stores and provides real-time feedback to store staff. The system is composed of a combination of cameras, audio collection devices, wireless communication connections, a server, smart devices, and a generative AI model.
[0763] Hardware and Software Configuration
[0764] Camera and audio collection device
[0765] Cameras are installed in various locations throughout the store and collect video and audio recordings of customers in real time. They have wireless communication capabilities and transmit video and audio data to a server. Specifically, store staff wear smart glasses (e.g., Google Glass or Vuzix M400) to collect customer interactions.
[0766] server
[0767] The server receives the collected video and audio data via wireless communication and analyzes it. This analysis involves identifying customers using biometric authentication technology and performing sentiment analysis using generative AI models (e.g., Affectiva, Microsoft Azure Cognitive Services). Furthermore, key scenes are identified based on the video and audio data, and editing is performed in real time.
[0768] Smart Devices
[0769] The smart device (smart glasses) displays the analysis results from the server, allowing store staff to instantly understand the customer's emotional state and respond appropriately.
[0770] Data processing and calculation
[0771] The server does the following:
[0772] 1. Receive video and audio data and identify customers using biometric technology.
[0773] 2. Use generative AI models to analyze customer sentiment in real time from video and audio data.
[0774] 3. Identify key scenes based on the customer's emotional state, edit in real time, and provide feedback to smart devices.
[0775] Specific examples
[0776] In practice, the following scenarios can be considered when using this system:
[0777] Situation 1: Dealing with customers in the fitting room
[0778] Example prompt sentence:
[0779] "If a staff member wearing glasses carefully observes a customer entering a fitting room and detects that the customer has a troubled expression."
[0780] What actually happens:
[0781] Smart glasses observe customers in fitting rooms.
[0782] The emotion engine analyzes the customer's facial expression to determine whether they are confused.
[0783] "The customer is in trouble" is displayed in real time on the smart glasses display.
[0784] Staff immediately went to the fitting room to provide the customer with appropriate assistance.
[0785] Situation 2: Serving customers at the counter
[0786] Example prompt sentence:
[0787] "When a staff member wearing glasses observes a customer who comes to a product counter, they detect that the customer has a satisfied expression."
[0788] What actually happens:
[0789] Smart glasses observe customers at the counter.
[0790] The emotion engine analyzes the customer's facial expression to determine whether they are satisfied.
[0791] "Customer is satisfied" is displayed in real time on the smart glasses display.
[0792] Staff will further enhance customer service, suggesting additional products and offering special services to customers.
[0793] This makes it possible to analyze customer sentiment in real time and provide appropriate support.
[0794] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0795] Step 1:
[0796] The cameras and audio collection devices begin to collect video and audio data in real time. This data captures how customers move around the store and what they say. The input is the video and audio captured by the cameras and microphones, and the output is the data transmitted wirelessly to the server.
[0797] Step 2:
[0798] The server receives and analyzes the data sent from the camera and audio collection device. Specifically, it uses a biometric authentication algorithm to identify customers from the video data. The input is the video and audio data sent from the camera and audio collection device, and the output is the ID information of the identified customer.
[0799] Step 3:
[0800] The server uses a generative AI model to analyze the customer's emotions from the received video and audio data, including facial expressions and tone of voice. The input is the identified customer's video and audio data, and the output is the customer's emotional state (e.g., interest, confusion, satisfaction, etc.).
[0801] Step 4:
[0802] The server identifies key moments in real time based on the customer's emotional state. Here, a generative AI model detects emotional peaks and changes and extracts specific events (e.g., excitement, confusion). The input is the customer's emotional state data, and the output is the timestamp and context of the identified key moments.
[0803] Step 5:
[0804] The server edits the video in real time based on the identified important scenes, cutting out unnecessary parts and emphasizing the important scenes. The input is the timestamp and context of the identified important scenes, and the output is the edited video clip.
[0805] Step 6:
[0806] The server sends the edited video clip to the smart device and provides feedback to the store staff in real time. The input is the edited video clip and the emotion analysis results, and the output is the feedback information displayed on the smart device's display.
[0807] Step 7:
[0808] Store staff wearing smart devices respond based on the feedback provided by the server. Specifically, if a customer's confusion is detected, the staff responds immediately and provides appropriate support. The input is the feedback information displayed on the smart device, and the output is the staff's immediate response.
[0809] Step 8:
[0810] The server stores all collected data and later analyzes it for long-term patterns and trends of customer behavior. The input is the entire video and audio data collected daily, and the output is regular reports and analysis results, providing information to support strategic decision-making.
[0811] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0812] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0813] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0814] [Third embodiment]
[0815] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0816] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0817] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0818] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0819] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0820] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0821] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0822] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0823] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0824] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0825] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0826] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0827] This invention is a system that uses multiple cameras to collect video and audio in real time, and a server analyzes and edits the video and audio. The specific program processing is explained below in natural language.
[0828] Camera Initial Settings
[0829] The terminal installs wireless LAN-compatible cameras in various locations in the venue. The cameras are placed on the ceiling, on tabletops, and around tables, and the terminal registers each camera's IP address, field of view, and installation location on the server.
[0830] Video collection and transmission
[0831] Cameras collect video and audio data from the entire venue in real time, and this data is sent to a server via wireless LAN.
[0832] Video and audio analysis
[0833] The server receives and analyzes the streaming data in real time. First, it uses facial recognition to identify guests based on group photos and video clips. Next, it uses AI to analyze the video and audio data and identify highlights based on emotional expressions and volume changes.
[0834] Video editing
[0835] The server then edits the identified highlights in real time, prioritizing important scenes such as the bride and groom's entrance, speeches, toasts, and cake cutting. The edited footage is then temporarily saved.
[0836] Attendee Video Integration
[0837] Users (attendees) upload video data taken with their smartphones or cameras to the device. The device then sends the uploaded data to the server, which receives and analyzes it. The received video data is integrated with existing highlight footage to create a more realistic video.
[0838] Generating the final video
[0839] The server generates the final video from all the video data, resulting in a complete video at the end of the event. The generated video file is then provided to users in an easily accessible format.
[0840] Re-editing function
[0841] The user sends a request to the server via their device to re-edit a specific scene. For example, if the user has a specific request, such as "I want more emphasis on the birthday cake," the server re-analyzes the video data and makes a new edit. The re-edited video is then provided to the user.
[0842] Specific examples
[0843] 1. Video recording of a wedding reception
[0844] The terminal places wireless LAN-compatible cameras at various locations throughout the venue and performs initial setup.
[0845] The camera captures the moment when the bride and groom enter, the guests give speeches, the toast, and the cake cutting in real time, and sends the video and audio data to the server.
[0846] The server analyzes the collected data and uses facial recognition to identify the bride and groom and important guests. Using AI, it extracts highlights based on emotional expressions and volume changes and edits them in real time.
[0847] Users (attendees) upload videos taken with their smartphones to the device and transfer them to the server.
[0848] The server then combines the uploaded footage with existing edited footage to create a richer video.
[0849] The server generates the final edited video and makes it ready to be served.
[0850] In this way, advanced video recording and editing can be achieved using AI and networks, without relying on the skills of cameramen or editors.
[0851] The processing flow will be explained below.
[0852] Step 1:
[0853] The terminal installs wireless LAN-compatible cameras in advance at various locations in the venue and registers information such as the camera's IP address, field of view, and installation location on the server.
[0854] Step 2:
[0855] The cameras begin collecting video and audio from the entire venue in real time and stream this data to a server via wireless LAN.
[0856] Step 3:
[0857] The server receives the streaming data sent from the camera in real time and begins analyzing the video and audio.
[0858] Step 4:
[0859] The server performs facial recognition based on group photos and data acquired in advance to identify guests.
[0860] Step 5:
[0861] The server uses AI to analyze video and audio, detecting emotional expressions and changes in volume to identify highlight scenes.
[0862] Step 6:
[0863] The server then prioritizes editing the identified highlight scenes, cutting out unnecessary parts to optimize scene transitions, and temporarily stores the edited footage.
[0864] Step 7:
[0865] Users (attendees) upload video data taken with their smartphones or cameras to the device.
[0866] Step 8:
[0867] The device sends the uploaded video data of attendees to the server, which receives and analyzes it.
[0868] Step 9:
[0869] The server then combines the footage uploaded by attendees with existing highlight footage to create a more immersive video.
[0870] Step 10:
[0871] The server generates the final video from all the video data and provides the video file in an easily accessible format for the user.
[0872] Step 11:
[0873] The user sends a request to the server via their device to re-edit a specific scene.
[0874] Step 12:
[0875] The server re-analyzes the video based on the re-editing request, performs a new edit, and provides the re-edited video to the user.
[0876] Example 1
[0877] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0878] Conventional video recording systems rely on manual processes for collecting, analyzing, and editing video and audio, resulting in limitations in real-time performance and accuracy. This makes efficient, high-quality video editing particularly difficult in situations where large amounts of data must be collected in a short period of time, such as events and meetings. It is also cumbersome for users to easily integrate and re-edit footage they have shot themselves. There is a demand for a system that can solve these issues and achieve efficient, high-precision video recording and editing.
[0879] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0880] In this invention, the server includes means for collecting video and audio from a camera in real time, means for transmitting the collected video and audio to the server via wireless communication, means for the server to identify guests using facial recognition, means for the server to analyze the video and audio using artificial intelligence and identify highlight scenes, means for editing in real time based on the highlight scenes identified by the server, means for uploading video data taken by attendees to the server using a terminal, means for integrating the uploaded video data of attendees with existing highlight videos, means for generating a final video using all the video data, means for re-editing based on a request to re-edit specific scenes, and means for outputting the generated final video and re-edited video, thereby enabling efficient and high-quality video recording and editing.
[0881] A "camera" is a photographic device that collects video and audio from within the venue in real time and transmits them to a server.
[0882] "Wireless communication" is a general term for technology that uses radio waves to send and receive data.
[0883] A "server" is a computer system that receives data from multiple cameras and devices via a network and analyzes, edits, saves, and outputs the data.
[0884] "Facial recognition" is a technology that analyzes faces contained in video data and identifies individuals.
[0885] "Artificial intelligence" is a data analysis technology that uses machine learning and deep learning to analyze video based on emotional expressions, volume changes, etc.
[0886] "Highlight scenes" refer to important moments or scenes with strong emotional expressions within video data.
[0887] "Real-time editing" refers to the process of collecting video data and simultaneously analyzing and editing it to create a format that can be output immediately.
[0888] An "attendee" is a person attending an event, meeting, etc., including anyone who captures video or audio on their own device.
[0889] "Terminals" are devices used by attendees to upload video data to the server. These devices include smartphones and tablets.
[0890] "Integration" refers to the process of analyzing multiple video data sets and editing and combining them into a series of flows.
[0891] The "final video" is the completed video edited based on all the collected video data.
[0892] "Re-editing" refers to the process of adding new edits to already edited video data based on a request from a user.
[0893] "Generation" refers to the process of creating the final video from the collected data.
[0894] "Output" refers to providing the generated Final Video or Re-Edited Video in a user-accessible format.
[0895] The present invention is a system that uses multiple cameras to collect video and audio in real time, and a server analyzes and edits the collected video and audio. Specifically, the system is implemented according to the following steps:
[0896] First, wireless communication-enabled cameras are installed in various locations throughout the venue. The cameras are placed on the ceiling, on tabletops, and around tables, and the IP address, field of view, and installation location of each camera are registered on the server by the terminal. At this time, a dedicated setting application is used to save the camera information in the server database.
[0897] Next, the camera collects video and audio from the entire venue in real time. The collected data is sent to the server via wireless communication as an H.264 video stream. The camera automatically starts transmitting data and provides a live feed to a URL specified by the server.
[0898] The server receives and analyzes the streaming data in real time. First, it performs facial recognition based on group photos and pre-collected video clips to identify guests. Specifically, the server calls an AI model to perform facial recognition and identify individual guests. Next, it uses the AI model to analyze the video and audio data and identify highlight scenes based on emotional expressions and volume changes. For example, it prioritizes the extraction of scenes featuring the bride and groom or scenes with moving speeches.
[0899] The server then edits the identified highlights in real time. Specifically, the server uses video editing software to edit the video. Important scenes, such as the bride and groom's entrance, the toast, and the cake cutting, are extracted as clips and placed on the timeline. The edited video is then temporarily saved in an appropriate format.
[0900] Users (attendees) upload video data they have taken with their own devices to their terminals. This data is then sent to the server using a dedicated upload application. The video data received by the server is synchronized and integrated with existing highlight footage. The server analyzes this data and inserts additional footage at the appropriate times to create a more immersive final video.
[0901] The server then combines all the video data to create the final video. Specifically, it combines each edited clip and organizes the overall timeline. The video format is a common format such as MP4 or MOV. The final video file is then uploaded to a cloud storage service, and a shared link is provided for users to easily access.
[0902] It is also possible for users to send requests to the server via their device to re-edit specific scenes. For example, if a user has a specific request, such as "I want more emphasis on the toast scene," the user enters the details in a dedicated form and sends it to the server. The server then analyzes the video data again and makes new edits. In this case, too, the re-editing is done using video editing software, and the re-edited video is provided to the user.
[0903] Specific examples
[0904] 1. Video recording of a wedding reception
[0905] Cameras with wireless communication capabilities will be placed around the venue, and initial settings will be performed using a dedicated application.
[0906] The camera captures the bride and groom's entrance, guest speeches, toasts, and cake cutting moments in real time as H.264 format video streams and sends them to the server.
[0907] The server analyzes the collected data and uses an artificial intelligence model to perform facial recognition to identify the bride and groom and important guests. AI is also used to extract highlights based on emotional expressions and volume changes, and the footage is then edited in real time using video editing software.
[0908] Users (attendees) upload the video they have taken with their smartphones to the device using a dedicated upload application, and then transfer it to the server.
[0909] The server analyzes the uploaded footage and combines it with existing edited footage to create a richer final video.
[0910] The server then combines all the clips, adds transitions and titles, and generates the final edited video, which is then made available to the user via cloud storage.
[0911] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0912] Step 1:
[0913] The terminal installs wireless communication-enabled cameras at various locations in the venue. First, a dedicated setting application is launched and the camera's IP address, field of view, and installation location are entered. This information is then sent to the server and registered in a database.
[0914] Input: Camera IP address, viewing angle, installation location
[0915] Output: Camera information stored in the server database
[0916] Step 2:
[0917] The camera collects video and audio from the venue in real time as an H.264 video stream, and transmits this data to a server via wireless communication. The camera automatically establishes a connection and provides a live feed to a specified URL.
[0918] Input: Video and audio collected by a camera
[0919] Output: H.264 video stream (sent to server)
[0920] Step 3:
[0921] The server receives and analyzes the streaming data in real time. It then activates a facial recognition system to identify guests based on pre-registered group photos and video clips. Specifically, the server invokes an artificial intelligence model to perform facial recognition.
[0922] Input: H.264 format video stream
[0923] Output: Identified guest data
[0924] Step 4:
[0925] The server uses an AI model to analyze video and audio data and identify highlights based on emotional expressions and changes in volume. For example, it can extract important scenes by detecting changes in the volume of smiles or applause.
[0926] Input: Video stream after facial recognition
[0927] Output: Identified highlight scene data
[0928] Step 5:
[0929] The server performs real-time editing based on the identified highlight scenes. Using Adobe Premiere Pro API and FFmpeg, important scenes are extracted as clips and placed on the timeline. The edited footage is temporarily saved in an appropriate format.
[0930] Input: Identified highlight scene data
[0931] Output: Edited video clip (temporarily saved)
[0932] Step 6:
[0933] Users (attendees) upload video data taken with their own devices to the terminal using a dedicated upload application, and the terminal then sends this data to the server.
[0934] Input: Video data taken by the user
[0935] Output: Video data uploaded to the server
[0936] Step 7:
[0937] The server analyzes the received video data and integrates it with the existing highlight footage, inserting additional footage into the existing data at the appropriate time to create a more immersive final video.
[0938] Input: Uploaded user video data
[0939] Output: Final merged video data
[0940] Step 8:
[0941] The server then combines all the video data to create the final video. Specifically, it combines the edited clips, adds transitions and titles, and creates a timeline. The final video is saved in a standard format (MP4 or MOV) and uploaded to cloud storage.
[0942] Input: Final merged video data
[0943] Output: Final video file saved in cloud storage
[0944] Step 9:
[0945] The user sends a request to re-edit a specific scene to the server via their device. They enter details into a request form and send it to the server. The server then analyzes the video data again and performs a new edit. This is done using video editing software.
[0946] Input: Request for re-edit
[0947] Output: Re-edited video
[0948] Step 10:
[0949] The server saves the re-edited video back to cloud storage and provides the user with a shared link, allowing them to easily access the newly edited video.
[0950] Input: Re-edited video
[0951] Output: Video Recut files saved in cloud storage and a share link
[0952] (Application example 1)
[0953] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0954] In traditional store operations, analyzing customer behavior, providing efficient customer service, implementing security measures, and implementing effective marketing often consumes a significant amount of time and cost. It can also be difficult to grasp the situation in real time and take prompt action. This creates challenges that prevent efficient store operations and improved customer experience.
[0955] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0956] In this invention, the server includes a means for analyzing video and audio data obtained by the camera in real time to identify people's movement paths and areas where they are staying, a means for notifying detected information in real time and encouraging proactive responses, and a means for generating a heat map of customer traffic and displaying it on an interface accessible by a manager, thereby making it possible to improve the efficiency of store operations and customer satisfaction.
[0957] A "camera" is a device used to collect video and audio and transmit it to a server via wireless communication.
[0958] A "server" is a computer system that analyzes collected video and audio data and performs various processes such as identification, editing, and notification.
[0959] "Machine learning" is a technology that analyzes video and audio data to automatically recognize and identify specific patterns and highlight scenes.
[0960] "Facial recognition" is a technology for recognizing a person's face from collected video data and identifying that person.
[0961] A "highlight scene" refers to a particularly important or interesting moment or action in video or audio data.
[0962] "Wireless communication" is a technology that uses wireless LAN to send and receive data.
[0963] "User" refers to an individual or organization that shoots video data and uploads the data to a server using a terminal.
[0964] A "terminal" is a device (such as a smartphone or tablet) that a user uses to upload video data to a server.
[0965] "Traffic lines" refer to the routes that show how people move around the store.
[0966] A "stay area" refers to an area where a person or customer stays in a particular location for a certain period of time.
[0967] "Real-time notifications" are notifications sent instantly to prompt necessary action by immediately processing collected information.
[0968] "Proactive response" means predicting the situation and providing appropriate measures and services immediately.
[0969] A "heat map" is a visual display of the popularity and frequency of use of a particular area, using color coding based on collected data.
[0970] This invention is a system that uses multiple cameras to collect video and audio in real time, and a server analyzes and edits the data. This system can be applied to store operations, customer service, security measures, and marketing optimization.
[0971] Hardware and software used
[0972] Hardware:
[0973] Wireless LAN camera: Collects video and audio in real time (e.g., a typical network camera).
[0974] Server: Analyzes video and audio data and performs various processes such as identification, editing, and notification (e.g., cloud servers such as AWS EC2).
[0975] Smartphone: A device used by users to shoot and upload video data (e.g., iPhone, Android device).
[0976] Robots: Serve customers and guide them in stores (e.g., Pepper robot).
[0977] software:
[0978] Video analytics AI: Analyzes video data to identify customer movement patterns and areas of congestion (e.g., OpenCV, TensorFlow).
[0979] Speech recognition AI: Converts voice data into text and recognizes questions or requests (e.g., Google Speech-to-Text API).
[0980] Database: Stores and manages collected data (e.g., AWS RDS).
[0981] Real-time communication: Providing immediate notifications (e.g., WebSockets).
[0982] Streaming services: Real-time delivery of video data (e.g., AWS Kinesis).
[0983] System Overview
[0984] 1. Camera Initial Setup:
[0985] The camera is connected to the wireless LAN, and the terminal registers the camera's IP address, field of view, and installation location on the server.
[0986] 2. Video Collection and Transmission:
[0987] The camera collects video and audio from the entire venue in real time and transmits it to a server via wireless communication.
[0988] 3. Video and audio analysis:
[0989] The server receives the streaming data in real time and uses machine learning to analyze the video and audio data, identifying people through facial recognition and pinpointing their movement patterns and areas of concentration.
[0990] 4. Real-time notifications and actions:
[0991] The detected information is sent in real time to the store clerk's smartphone or the robot, encouraging proactive response.
[0992] 5. Heatmap and report generation:
[0993] Customer traffic heat maps are generated and displayed in the admin interface, and regular reports are generated based on the collected and analyzed data.
[0994] 6. Re-editing and final video creation:
[0995] The server re-edits specific scenes as needed and generates and serves the final video in response to the user's request.
[0996] Specific examples
[0997] Customer behavior analysis:
[0998] Cameras track customers' movements within the store and measure the time they spend in front of specific product areas, and the server then reports back to them the most popular products in the areas where customers spend the most time.
[0999] Prompt Sentence Examples
[1000] Customer behavior analysis:
[1001] - "Please identify customer movement patterns and areas where customers stay using this camera video data."
[1002] - "Please recognize the customer's question from this voice data."
[1003] This system will improve the efficiency of store operations and customer satisfaction, allowing for real-time situation assessment and rapid response.
[1004] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1005] Step 1:
[1006] Camera Defaults:
[1007] The terminal installs wireless LAN-enabled cameras in various areas of the store. The terminal registers the camera's IP address, field of view, and installation location on the server. The terminal receives the camera's physical installation location and setting information as input, and stores this information in a database as output.
[1008] Step 2:
[1009] Video collection and transmission:
[1010] The cameras collect video and audio from the entire venue in real time. This data is sent to the server via wireless communication. The system receives video and audio data from the cameras as input and sends it to the server via wireless communication. The output is treated as streaming data.
[1011] Step 3:
[1012] Video and audio analysis:
[1013] The server receives the streaming data in real time and analyzes it using video analysis AI and voice recognition AI. This allows it to identify customer movement patterns, areas where customers stay, personal identification using facial recognition, and the details of their questions and requests. It receives streaming data as input and stores the identified information (movement patterns, areas where customers stay, personal identification results, and questions) as output in a database.
[1014] Step 4:
[1015] Real-time notifications and actions:
[1016] The server notifies the detected information in real time to the store clerk's smartphone or the in-store robot, allowing the store clerk or robot to proactively respond to customers. The server receives the analysis results as input and sends notification data as output via WebSocket.
[1017] Step 5:
[1018] Generate a heatmap:
[1019] The server generates a heat map of customer traffic within the store based on data on customer movement and lingering areas. The generated heat map is displayed on an interface that can be accessed by administrators. The server receives data on movement and lingering areas as input, generates a heat map as output, and displays it on the interface.
[1020] Step 6:
[1021] Report Generation:
[1022] The server periodically generates reports based on the collected and analyzed data. The reports include customer movement patterns, areas where customers stay, questions, etc. It receives various data stored in the database as input, generates reports as output, and sends them to the administrator.
[1023] Step 7:
[1024] Re-editing and producing the final video:
[1025] The server re-edits specific scenes as needed and generates and provides the final video in response to the user's request. It receives the user's request and video data as input, generates the re-edited video as output, and provides it to the user.
[1026] In this way, real-time data collection and analysis, immediate notification and response, and visualization and reporting of analysis results will lead to more efficient store operations and improved customer satisfaction.
[1027] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1028] This invention combines a system that uses multiple cameras to collect and analyze video and audio in real time with an emotion engine. Below, we will explain the specific program processing of this system in natural language.
[1029] Camera Initial Settings
[1030] The terminal places wireless LAN-compatible cameras in various locations in the venue. The cameras are installed on the ceiling, on tabletops, and around tables, and registers information such as the camera's IP address, field of view, and installation location on the server.
[1031] Video collection and transmission
[1032] The camera begins collecting video and audio from the entire venue in real time and transmits it to a server via wireless LAN.
[1033] Video and audio analysis
[1034] The server receives the streaming data sent from the camera in real time, analyzes the video and audio, and primarily performs facial recognition to identify guests based on group photos and data acquired in advance.
[1035] Incorporating an emotion engine
[1036] The server uses an emotion engine to analyze the user's emotional state in real time based on the collected video and audio data. Based on this emotional state, highlight scenes are identified according to emotional expressions and volume changes in the video. For example, moments of cheering or crying are automatically extracted.
[1037] Video editing
[1038] The server edits the video in real time based on the identified highlight scenes, cutting out unnecessary parts and optimizing scene transitions. It also appropriately edits scenes that should be emphasized according to the emotional state identified by the emotion engine.
[1039] Attendee Video Integration
[1040] Users (attendees) upload video data taken with their smartphones or cameras to their devices. The devices then send the uploaded data to the server, which receives and analyzes it. The uploaded video data is integrated into existing highlight footage based on the analysis results of the emotion engine, creating a more moving and realistic video.
[1041] Generating the final video
[1042] The server generates the final video from all the video data, and the resulting video file is provided to the user in an easily accessible format.
[1043] Re-editing function
[1044] The user sends a request to the server via their device to re-edit a specific scene. For example, a specific request such as "I want more emphasis on the bride and groom's reaction" can be accepted. The server re-analyzes the video data based on the re-editing request and performs a new edit. The re-edited video is then provided to the user.
[1045] Specific examples
[1046] 1. Video recording of a wedding reception
[1047] The terminal places wireless LAN-compatible cameras at various locations throughout the venue and performs initial setup.
[1048] The camera captures the moment when the bride and groom appear, the guests' speeches, the toast, the cake cutting, and the expressions of the many guests in real time, and sends the video and audio data to the server.
[1049] The server analyzes the transmitted data and identifies the bride and groom and important guests using facial recognition, analyzes their emotional state using an emotion engine, and identifies highlight scenes based on emotional expressions and volume changes.
[1050] The server edits the video in real time based on the identified highlights. For example, the emotion engine detects the bride's tears of joy and emphasizes these scenes.
[1051] Users (attendees) upload videos they have taken with their smartphones, which are then received and analyzed by the server. The uploaded video is then integrated with existing edited footage based on the analysis results of the emotion engine, allowing for a richer reproduction of the bride and groom's reactions and the guests' emotional moments.
[1052] The server generates the final edited video and makes it ready to be served.
[1053] In this way, a system that combines an emotion engine enables high-quality video recording and editing that reflects the user's emotional state.
[1054] The processing flow will be explained below.
[1055] Step 1:
[1056] The terminal installs wireless LAN-compatible cameras in advance at various locations in the venue and registers information such as the camera's IP address, field of view, and installation location on the server.
[1057] Step 2:
[1058] The cameras begin collecting video and audio from the entire venue in real time and streaming it to a server via Wi-Fi.
[1059] Step 3:
[1060] The server receives the streaming data sent from the camera in real time and analyzes the video and audio.
[1061] Step 4:
[1062] The server performs facial recognition based on group photos and data acquired in advance to identify guests.
[1063] Step 5:
[1064] The server uses AI to analyze video and audio data, detecting emotional expressions and changes in volume to identify highlight scenes.
[1065] Step 6:
[1066] The server uses an emotion engine to analyze the user's emotional state in real time, and identifies important scenes in the video based on the emotional state.
[1067] Step 7:
[1068] The server then prioritizes editing the identified highlights and optimizes scene transitions by cutting out unnecessary parts, for example, by emphasizing touching moments between the bride and groom or scenes that elicit cheers of joy.
[1069] Step 8:
[1070] Users (attendees) upload video data taken with their smartphones or cameras to the device.
[1071] Step 9:
[1072] The device sends the uploaded video data of attendees to the server, which receives and analyzes it.
[1073] Step 10:
[1074] The server then integrates the uploaded attendee video data with existing highlight footage based on the analysis results of the emotion engine, creating a more moving and realistic video.
[1075] Step 11:
[1076] The server generates the final video from all the video data and provides the video file in an easily accessible format for the user.
[1077] Step 12:
[1078] The user sends a request to the server via their device to re-edit a specific scene, such as "I want more emphasis on the reaction of the bride and groom."
[1079] Step 13:
[1080] The server re-analyzes the video data based on the re-editing request and performs a new edit. The re-edited video is then provided to the user.
[1081] Example 2
[1082] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1083] Modern events require efficient recording and editing of many important moments. However, doing so manually is a time-consuming and labor-intensive task. Furthermore, integrating footage shot separately by attendees is tedious and does not necessarily produce natural-looking results. Furthermore, capturing emotional expressions and volume changes in the footage and identifying moving moments and important scenes is a difficult challenge. Therefore, there is a need for a system that can automatically collect, analyze, and edit footage in real time, and efficiently integrate additional footage from users.
[1084] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1085] In this invention, the server includes means for collecting video and audio from a camera in real time, means for transmitting the collected video and audio to the server via wireless communication, means for the server to perform individual identification using facial recognition technology, means for analyzing the video and audio using machine learning to identify important scenes, means for editing in real time based on the important scenes identified by the server, means for uploading video data shot by a user to the server using a terminal, means for integrating the uploaded user's video data with existing important scene videos, means for generating a final video using all the video data, means for re-editing based on a request to re-edit a specific scene, and means for outputting the generated final video and re-edited video. This makes it possible to automatically and efficiently collect, analyze, and edit video, and naturally integrate additional user video.
[1086] A "camera" is a device for collecting video and audio and capable of transmitting data to a server via a wireless communication network.
[1087] "Server" is a central processing unit for receiving, analyzing, and editing collected video and audio data.
[1088] "Facial recognition technology" refers to algorithms and techniques for identifying people from video data.
[1089] "Machine learning" is a technology that analyzes large amounts of data to learn patterns and characteristics and make predictions and distinctions.
[1090] An "important scene" refers to a particularly noteworthy or moving moment in a video, and is identified by emotional expressions or changes in volume.
[1091] "Real-time editing" is the process of editing scenes that are deemed important at the moment as video data is collected.
[1092] "User-recorded video data" refers to video files that event attendees record themselves using their smartphones or cameras and later upload to the system.
[1093] "Uploading" refers to the act of sending video data taken by a user to a server.
[1094] "Merge" is the process of combining uploaded footage with existing footage and key scenes.
[1095] The "final video" is a video file in the final format after all video data has been edited and compiled.
[1096] "Re-editing" is a process in which already edited video data is re-edited based on a specific user request.
[1097] "Generated final video or re-edited video" includes final version video files automatically generated by the system and re-edited video files based on user requests.
[1098] A "wireless communication enabled camera" is a camera that has the ability to transmit data using a wireless communication network.
[1099] A "communications network" is an infrastructure for transmitting and receiving data between multiple devices using wireless LAN or other wireless communication protocols.
[1100] "Emotional expression" refers to the psychological state and emotional expression obtained by analyzing facial expressions in the video.
[1101] "Volume change" refers to an increase or decrease in volume detected by analyzing the audio data in the video.
[1102] The present invention combines a system that utilizes multiple wireless communication enabled cameras to collect and analyze video and audio in real time with an emotion engine. Specific embodiments for implementing this system are described below.
[1103] First, wireless communication-enabled cameras are installed in the venue. The device (e.g., laptop or tablet PC) sets the camera's IP address, field of view, and installation location, and registers this data on the server. The cameras are installed on the ceiling, on tabletops, around tables, etc., and are positioned so that they cover the entire venue. At this point, the cameras and server are connected using a wireless communication network (such as wireless LAN).
[1104] Next, the camera starts collecting video and audio in real time. The camera also collects audio data using its built-in microphone and transmits this data wirelessly to the server. The server uses a high-performance CPU and GPU to receive and analyze the streaming data sent from the camera in real time.
[1105] The server first uses facial recognition technology (e.g., OpenCV or dlib) to identify individuals from the collected video data, then uses an algorithm to analyze the audio data, detecting specific keywords and volume changes. The identified data is then stored in a database.
[1106] Furthermore, the server uses an emotion engine (e.g., Emotion API or Affectiva) to analyze the user's emotional state from video and audio data. The emotion engine analyzes facial expressions and audio tones for each frame to identify emotional states such as smiling or crying faces. This allows highlight scenes (important scenes) to be automatically identified.
[1107] Based on the identified highlight scenes, the server edits the video in real time. For editing, editing software (e.g., FFmpeg) is used to cut out unnecessary parts and ensure smooth scene transitions. Scenes that should be emphasized are appropriately edited according to emotional expressions and volume changes.
[1108] In addition, users (attendees) can upload video data taken with their smartphones or cameras to their devices, which then send the uploaded data to the server. The server receives the data, analyzes it using an emotion engine, and integrates it into existing highlight footage. This creates a more moving and realistic video.
[1109] Finally, the server generates the final video from all the video data, and the completed video file is uploaded to cloud storage and provided to users in a format that is easy to access (e.g., MP4, AVI).
[1110] Furthermore, users can send requests to the server via their devices to re-edit specific scenes. For example, a specific request could be made to "emphasize the bride and groom's reactions more." The server then re-analyzes the original editing data based on the re-editing request and performs a new edit. The re-edited video is also uploaded to cloud storage and provided to the user.
[1111] Specific examples
[1112] 1. Video recording of a wedding reception
[1113] The terminals are used to place wireless communication-enabled cameras at various locations in the venue, and the initial settings are performed. The camera's IP address, field of view, and installation location are registered on the server.
[1114] The camera captures the moment when the bride and groom appear, the guests' speeches, the toast, the cake cutting, and the expressions of the many guests in real time, and sends the video and audio data to the server.
[1115] The server analyzes the transmitted data and uses facial recognition to identify the bride and groom and important guests, and an emotion engine to analyze their emotional state and identify highlight scenes based on emotional expressions and volume changes.
[1116] The server edits the video in real time based on the identified highlights. For example, the emotion engine detects the bride's tears of joy and emphasizes these scenes.
[1117] Users (attendees) upload videos they have taken with their smartphones, which are then received and analyzed by the server. The uploaded video is then integrated with existing edited footage based on the analysis results of the emotion engine, allowing for a richer reproduction of the bride and groom's reactions and the guests' emotional moments.
[1118] The server generates the final edited video and makes it ready to be served. The video file is uploaded to cloud storage in MP4 format and a download link is provided to the user.
[1119] Example prompts to input to the generative AI model
[1120] "I want to create a video that selects and edits the most touching scenes from a wedding reception."
[1121] "Create a highlight video that highlights the bride and groom's reactions and the guests' joyful moments."
[1122] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1123] Step 1: Initial camera setup
[1124] The terminal places wireless communication-enabled cameras in the venue. Specifically, it sets the camera's IP address, field of view, and installation location information, and registers this data on the server. For example, the terminal checks the network connection of each camera from a settings screen and enters the setting information. The input is the camera's physical installation location, field of view, and IP address. The output is the setting information stored in the server's database.
[1125] Step 2: Collect and transmit video and audio
[1126] The camera starts collecting video and audio in real time and sends them to the server via wireless communication. Specifically, the camera captures frames every second and collects audio data. The input is the video and audio acquired from the camera's sensor. The output is data sent to the server in streaming format.
[1127] Step 3: Video and audio analysis
[1128] The server receives and analyzes streaming data sent from the camera in real time. Individual identification is performed using facial recognition technology (e.g., OpenCV or dlib). The input is video data sent from the camera. Specifically, the captured frames are run through a facial recognition algorithm to detect the position of the face and its features. The output is information about the identified person and feature data. Audio data is also analyzed to detect specific keywords and volume changes. The input is audio data sent from the camera. The output is keyword detection information and volume change information as the analysis results.
[1129] Step 4: Incorporating the Emotion Engine
[1130] The server uses an emotion engine (e.g., Emotion API or Affectiva) to analyze the user's emotional state from video and audio data. The input is the video and audio feature data analyzed in the previous step. Specifically, the emotion engine analyzes facial expressions and audio tone for each frame to identify emotional states such as laughter or tears. The output is data related to the emotional state.
[1131] Step 5: Edit your footage
[1132] The server edits the video in real time based on the highlight scenes identified by the emotion engine. Specifically, it uses editing software (e.g., FFmpeg) to cut out unnecessary parts and smooth scene transitions. The input is the identified highlight scene information and the original video data. The output is the edited video data.
[1133] Step 6: Integrate attendee videos
[1134] Users (attendees) upload video data taken with their smartphones or cameras to their terminals. The terminals then send the uploaded data to the server. Specifically, the terminals manage the process of selecting video files and transferring them to the server. The input is the video data uploaded from the user's device. The output is additional video data stored on the server. The server analyzes the data using an emotion engine and integrates it into the existing highlight video. The input is the newly uploaded video data and the existing highlight video data. The output is the integrated video data.
[1135] Step 7: Generate the final video
[1136] The server generates the final video from all the video data. Specifically, editing software is used to combine all the scenes and create the final video file. The input is the merged video data. The output is the final video file (e.g., MP4 format). The generated file is uploaded to cloud storage so that users can access it.
[1137] Step 8: Re-edit function
[1138] The user sends a request to the server via their device to re-edit a specific scene. For example, a request may be made to "emphasize the bride and groom's reaction more." The server re-analyzes the original video data based on the re-editing request and performs a new edit. Specifically, the server re-processes the original video data according to the re-editing instructions and edits it to emphasize the necessary parts. The input is the re-editing request and the original video data. The output is a re-edited video file, which is also uploaded to cloud storage and a download link is provided to the user.
[1139] (Application example 2)
[1140] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1141] There is a need for a system that can analyze customer emotions in real time and provide appropriate feedback to store staff to provide more effective customer service and improve customer satisfaction. However, existing systems lack the means to analyze customer emotions and behavior in real time, making it difficult for staff to respond immediately.
[1142] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting video and audio from a camera in real time, means for transmitting the collected video and audio to the server via wireless communication, means for the server to identify people using biometric authentication, means for the server to analyze the video and audio using a generative AI model and identify key scenes, means for editing in real time based on the key scenes identified by the server, means for uploading video data taken by participants to the server using a terminal, means for integrating the uploaded video data of participants with existing key scene videos, means for generating a final video using all video data, means for re-editing based on a request to re-edit a specific scene, means for outputting the generated final video and re-edited video, and means for the smart device to analyze customer emotions in real time and provide feedback to staff. This makes it possible to analyze customer emotions in real time within the store and provide prompt and appropriate support.
[1143] A "camera" is a device that collects video and audio and converts it into electrical signals.
[1144] "Wireless communication" is a means of sending and receiving data without using cables, and is a technology that uses radio signals.
[1145] A "server" is a computer system that processes and stores data on a network and provides various services.
[1146] "Biometric authentication" is a technology that identifies individuals using biological characteristics such as their face, fingerprints, and irises.
[1147] A "generative AI model" is an algorithm or program that uses artificial intelligence to analyze data and generate new information or results.
[1148] "Video" is visual data generated by optical means and collected by a camera or other imaging device.
[1149] "Sound" refers to a sound signal transmitted by vibrations in the air, and is collected by a sound collection device such as a microphone.
[1150] An "important scene" refers to a scene or moment in video or audio data that is particularly noteworthy.
[1151] "Real-time" refers to data processing and analysis without delay, providing results almost instantly.
[1152] A "terminal" is a device that connects to a network and transmits, receives, and processes data.
[1153] "Integration" refers to combining multiple data or elements into one.
[1154] A "participant" is an individual or group who is actively involved in an event, activity, or project.
[1155] A "smart device" is an advanced electronic device that has communication capabilities and processes and presents information.
[1156] This invention describes a system that analyzes customer sentiment in brick-and-mortar stores and provides real-time feedback to store staff. The system is composed of a combination of cameras, audio collection devices, wireless communication connections, a server, smart devices, and a generative AI model.
[1157] Hardware and Software Configuration
[1158] Camera and audio collection device
[1159] Cameras are installed in various locations throughout the store and collect video and audio recordings of customers in real time. They have wireless communication capabilities and transmit video and audio data to a server. Specifically, store staff wear smart glasses (e.g., Google Glass or Vuzix M400) to collect customer interactions.
[1160] server
[1161] The server receives the collected video and audio data via wireless communication and analyzes it. This analysis involves identifying customers using biometric authentication technology and performing sentiment analysis using generative AI models (e.g., Affectiva, Microsoft Azure Cognitive Services). Furthermore, key scenes are identified based on the video and audio data, and editing is performed in real time.
[1162] Smart Devices
[1163] The smart device (smart glasses) displays the analysis results from the server, allowing store staff to instantly understand the customer's emotional state and respond appropriately.
[1164] Data processing and calculation
[1165] The server does the following:
[1166] 1. Receive video and audio data and identify customers using biometric technology.
[1167] 2. Use generative AI models to analyze customer sentiment in real time from video and audio data.
[1168] 3. Identify key scenes based on the customer's emotional state, edit in real time, and provide feedback to smart devices.
[1169] Specific examples
[1170] In practice, the following scenarios can be considered when using this system:
[1171] Situation 1: Dealing with customers in the fitting room
[1172] Example prompt sentence:
[1173] "If a staff member wearing glasses carefully observes a customer entering a fitting room and detects that the customer has a troubled expression."
[1174] What actually happens:
[1175] Smart glasses observe customers in fitting rooms.
[1176] The emotion engine analyzes the customer's facial expression to determine whether they are confused.
[1177] "The customer is in trouble" is displayed in real time on the smart glasses display.
[1178] Staff immediately went to the fitting room to provide the customer with appropriate assistance.
[1179] Situation 2: Serving customers at the counter
[1180] Example prompt sentence:
[1181] "When a staff member wearing glasses observes a customer who comes to a product counter, they detect that the customer has a satisfied expression."
[1182] What actually happens:
[1183] Smart glasses observe customers at the counter.
[1184] The emotion engine analyzes the customer's facial expression to determine whether they are satisfied.
[1185] "Customer is satisfied" is displayed in real time on the smart glasses display.
[1186] Staff will further enhance customer service, suggesting additional products and offering special services to customers.
[1187] This makes it possible to analyze customer sentiment in real time and provide appropriate support.
[1188] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1189] Step 1:
[1190] The cameras and audio collection devices begin to collect video and audio data in real time. This data captures how customers move around the store and what they say. The input is the video and audio captured by the cameras and microphones, and the output is the data transmitted wirelessly to the server.
[1191] Step 2:
[1192] The server receives and analyzes the data sent from the camera and audio collection device. Specifically, it uses a biometric authentication algorithm to identify customers from the video data. The input is the video and audio data sent from the camera and audio collection device, and the output is the ID information of the identified customer.
[1193] Step 3:
[1194] The server uses a generative AI model to analyze the customer's emotions from the received video and audio data, including facial expressions and tone of voice. The input is the identified customer's video and audio data, and the output is the customer's emotional state (e.g., interest, confusion, satisfaction, etc.).
[1195] Step 4:
[1196] The server identifies key moments in real time based on the customer's emotional state. Here, a generative AI model detects emotional peaks and changes and extracts specific events (e.g., excitement, confusion). The input is the customer's emotional state data, and the output is the timestamp and context of the identified key moments.
[1197] Step 5:
[1198] The server edits the video in real time based on the identified important scenes, cutting out unnecessary parts and emphasizing the important scenes. The input is the timestamp and context of the identified important scenes, and the output is the edited video clip.
[1199] Step 6:
[1200] The server sends the edited video clip to the smart device and provides feedback to the store staff in real time. The input is the edited video clip and the emotion analysis results, and the output is the feedback information displayed on the smart device's display.
[1201] Step 7:
[1202] Store staff wearing smart devices respond based on the feedback provided by the server. Specifically, if a customer's confusion is detected, the staff responds immediately and provides appropriate support. The input is the feedback information displayed on the smart device, and the output is the staff's immediate response.
[1203] Step 8:
[1204] The server stores all collected data and later analyzes it for long-term patterns and trends of customer behavior. The input is the entire video and audio data collected daily, and the output is regular reports and analysis results, providing information to support strategic decision-making.
[1205] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1206] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1207] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1208] [Fourth embodiment]
[1209] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1210] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1211] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1212] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1213] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1214] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1215] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1216] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1217] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1218] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1219] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1220] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1221] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1222] This invention is a system that uses multiple cameras to collect video and audio in real time, and a server analyzes and edits the video and audio. The specific program processing is explained below in natural language.
[1223] Camera Initial Settings
[1224] The terminal installs wireless LAN-compatible cameras in various locations in the venue. The cameras are placed on the ceiling, on tabletops, and around tables, and the terminal registers each camera's IP address, field of view, and installation location on the server.
[1225] Video collection and transmission
[1226] Cameras collect video and audio data from the entire venue in real time, and this data is sent to a server via wireless LAN.
[1227] Video and audio analysis
[1228] The server receives and analyzes the streaming data in real time. First, it uses facial recognition to identify guests based on group photos and video clips. Next, it uses AI to analyze the video and audio data and identify highlights based on emotional expressions and volume changes.
[1229] Video editing
[1230] The server then edits the identified highlights in real time, prioritizing important scenes such as the bride and groom's entrance, speeches, toasts, and cake cutting. The edited footage is then temporarily saved.
[1231] Attendee Video Integration
[1232] Users (attendees) upload video data taken with their smartphones or cameras to the device. The device then sends the uploaded data to the server, which receives and analyzes it. The received video data is integrated with existing highlight footage to create a more realistic video.
[1233] Generating the final video
[1234] The server generates the final video from all the video data, resulting in a complete video at the end of the event. The generated video file is then provided to users in an easily accessible format.
[1235] Re-editing function
[1236] The user sends a request to the server via their device to re-edit a specific scene. For example, if the user has a specific request, such as "I want more emphasis on the birthday cake," the server re-analyzes the video data and makes a new edit. The re-edited video is then provided to the user.
[1237] Specific examples
[1238] 1. Video recording of a wedding reception
[1239] The terminal places wireless LAN-compatible cameras at various locations throughout the venue and performs initial setup.
[1240] The camera captures the moment when the bride and groom enter, the guests give speeches, the toast, and the cake cutting in real time, and sends the video and audio data to the server.
[1241] The server analyzes the collected data and uses facial recognition to identify the bride and groom and important guests. Using AI, it extracts highlights based on emotional expressions and volume changes and edits them in real time.
[1242] Users (attendees) upload videos taken with their smartphones to the device and transfer them to the server.
[1243] The server then combines the uploaded footage with existing edited footage to create a richer video.
[1244] The server generates the final edited video and makes it ready to be served.
[1245] In this way, advanced video recording and editing can be achieved using AI and networks, without relying on the skills of cameramen or editors.
[1246] The processing flow will be explained below.
[1247] Step 1:
[1248] The terminal installs wireless LAN-compatible cameras in advance at various locations in the venue and registers information such as the camera's IP address, field of view, and installation location on the server.
[1249] Step 2:
[1250] The cameras begin collecting video and audio from the entire venue in real time and stream this data to a server via wireless LAN.
[1251] Step 3:
[1252] The server receives the streaming data sent from the camera in real time and begins analyzing the video and audio.
[1253] Step 4:
[1254] The server performs facial recognition based on group photos and data acquired in advance to identify guests.
[1255] Step 5:
[1256] The server uses AI to analyze video and audio, detecting emotional expressions and changes in volume to identify highlight scenes.
[1257] Step 6:
[1258] The server then prioritizes editing the identified highlight scenes, cutting out unnecessary parts to optimize scene transitions, and temporarily stores the edited footage.
[1259] Step 7:
[1260] Users (attendees) upload video data taken with their smartphones or cameras to the device.
[1261] Step 8:
[1262] The device sends the uploaded video data of attendees to the server, which receives and analyzes it.
[1263] Step 9:
[1264] The server then combines the footage uploaded by attendees with existing highlight footage to create a more immersive video.
[1265] Step 10:
[1266] The server generates the final video from all the video data and provides the video file in an easily accessible format for the user.
[1267] Step 11:
[1268] The user sends a request to the server via their device to re-edit a specific scene.
[1269] Step 12:
[1270] The server re-analyzes the video based on the re-editing request, performs a new edit, and provides the re-edited video to the user.
[1271] Example 1
[1272] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1273] Conventional video recording systems rely on manual processes for collecting, analyzing, and editing video and audio, resulting in limitations in real-time performance and accuracy. This makes efficient, high-quality video editing particularly difficult in situations where large amounts of data must be collected in a short period of time, such as events and meetings. It is also cumbersome for users to easily integrate and re-edit footage they have shot themselves. There is a demand for a system that can solve these issues and achieve efficient, high-precision video recording and editing.
[1274] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1275] In this invention, the server includes means for collecting video and audio from a camera in real time, means for transmitting the collected video and audio to the server via wireless communication, means for the server to identify guests using facial recognition, means for the server to analyze the video and audio using artificial intelligence and identify highlight scenes, means for editing in real time based on the highlight scenes identified by the server, means for uploading video data taken by attendees to the server using a terminal, means for integrating the uploaded video data of attendees with existing highlight videos, means for generating a final video using all the video data, means for re-editing based on a request to re-edit specific scenes, and means for outputting the generated final video and re-edited video, thereby enabling efficient and high-quality video recording and editing.
[1276] A "camera" is a photographic device that collects video and audio from within the venue in real time and transmits them to a server.
[1277] "Wireless communication" is a general term for technology that uses radio waves to send and receive data.
[1278] A "server" is a computer system that receives data from multiple cameras and devices via a network and analyzes, edits, saves, and outputs the data.
[1279] "Facial recognition" is a technology that analyzes faces contained in video data and identifies individuals.
[1280] "Artificial intelligence" is a data analysis technology that uses machine learning and deep learning to analyze video based on emotional expressions, volume changes, etc.
[1281] "Highlight scenes" refer to important moments or scenes with strong emotional expressions within video data.
[1282] "Real-time editing" refers to the process of collecting video data and simultaneously analyzing and editing it to create a format that can be output immediately.
[1283] An "attendee" is a person attending an event, meeting, etc., including anyone who captures video or audio on their own device.
[1284] "Terminals" are devices used by attendees to upload video data to the server. These devices include smartphones and tablets.
[1285] "Integration" refers to the process of analyzing multiple video data sets and editing and combining them into a series of flows.
[1286] The "final video" is the completed video edited based on all the collected video data.
[1287] "Re-editing" refers to the process of adding new edits to already edited video data based on a request from a user.
[1288] "Generation" refers to the process of creating the final video from the collected data.
[1289] "Output" refers to providing the generated Final Video or Re-Edited Video in a user-accessible format.
[1290] The present invention is a system that uses multiple cameras to collect video and audio in real time, and a server analyzes and edits the collected video and audio. Specifically, the system is implemented according to the following steps:
[1291] First, wireless communication-enabled cameras are installed in various locations throughout the venue. The cameras are placed on the ceiling, on tabletops, and around tables, and the IP address, field of view, and installation location of each camera are registered on the server by the terminal. At this time, a dedicated setting application is used to save the camera information in the server database.
[1292] Next, the camera collects video and audio from the entire venue in real time. The collected data is sent to the server via wireless communication as an H.264 video stream. The camera automatically starts transmitting data and provides a live feed to a URL specified by the server.
[1293] The server receives and analyzes the streaming data in real time. First, it performs facial recognition based on group photos and pre-collected video clips to identify guests. Specifically, the server calls an AI model to perform facial recognition and identify individual guests. Next, it uses the AI model to analyze the video and audio data and identify highlight scenes based on emotional expressions and volume changes. For example, it prioritizes the extraction of scenes featuring the bride and groom or scenes with moving speeches.
[1294] The server then edits the identified highlights in real time. Specifically, the server uses video editing software to edit the video. Important scenes, such as the bride and groom's entrance, the toast, and the cake cutting, are extracted as clips and placed on the timeline. The edited video is then temporarily saved in an appropriate format.
[1295] Users (attendees) upload video data they have taken with their own devices to their terminals. This data is then sent to the server using a dedicated upload application. The video data received by the server is synchronized and integrated with existing highlight footage. The server analyzes this data and inserts additional footage at the appropriate times to create a more immersive final video.
[1296] The server then combines all the video data to create the final video. Specifically, it combines each edited clip and organizes the overall timeline. The video format is a common format such as MP4 or MOV. The final video file is then uploaded to a cloud storage service, and a shared link is provided for users to easily access.
[1297] It is also possible for users to send requests to the server via their device to re-edit specific scenes. For example, if a user has a specific request, such as "I want more emphasis on the toast scene," the user enters the details in a dedicated form and sends it to the server. The server then analyzes the video data again and makes new edits. In this case, too, the re-editing is done using video editing software, and the re-edited video is provided to the user.
[1298] Specific examples
[1299] 1. Video recording of a wedding reception
[1300] Cameras with wireless communication capabilities will be placed around the venue, and initial settings will be performed using a dedicated application.
[1301] The camera captures the bride and groom's entrance, guest speeches, toasts, and cake cutting moments in real time as H.264 format video streams and sends them to the server.
[1302] The server analyzes the collected data and uses an artificial intelligence model to perform facial recognition to identify the bride and groom and important guests. AI is also used to extract highlights based on emotional expressions and volume changes, and the footage is then edited in real time using video editing software.
[1303] Users (attendees) upload the video they have taken with their smartphones to the device using a dedicated upload application, and then transfer it to the server.
[1304] The server analyzes the uploaded footage and combines it with existing edited footage to create a richer final video.
[1305] The server then combines all the clips, adds transitions and titles, and generates the final edited video, which is then made available to the user via cloud storage.
[1306] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1307] Step 1:
[1308] The terminal installs wireless communication-enabled cameras at various locations in the venue. First, a dedicated setting application is launched and the camera's IP address, field of view, and installation location are entered. This information is then sent to the server and registered in a database.
[1309] Input: Camera IP address, viewing angle, installation location
[1310] Output: Camera information stored in the server database
[1311] Step 2:
[1312] The camera collects video and audio from the venue in real time as an H.264 video stream, and transmits this data to a server via wireless communication. The camera automatically establishes a connection and provides a live feed to a specified URL.
[1313] Input: Video and audio collected by a camera
[1314] Output: H.264 video stream (sent to server)
[1315] Step 3:
[1316] The server receives and analyzes the streaming data in real time. It then activates a facial recognition system to identify guests based on pre-registered group photos and video clips. Specifically, the server invokes an artificial intelligence model to perform facial recognition.
[1317] Input: H.264 format video stream
[1318] Output: Identified guest data
[1319] Step 4:
[1320] The server uses an AI model to analyze video and audio data and identify highlights based on emotional expressions and changes in volume. For example, it can extract important scenes by detecting changes in the volume of smiles or applause.
[1321] Input: Video stream after facial recognition
[1322] Output: Identified highlight scene data
[1323] Step 5:
[1324] The server performs real-time editing based on the identified highlight scenes. Using Adobe Premiere Pro API and FFmpeg, important scenes are extracted as clips and placed on the timeline. The edited footage is temporarily saved in an appropriate format.
[1325] Input: Identified highlight scene data
[1326] Output: Edited video clip (temporarily saved)
[1327] Step 6:
[1328] Users (attendees) upload video data taken with their own devices to the terminal using a dedicated upload application, and the terminal then sends this data to the server.
[1329] Input: Video data taken by the user
[1330] Output: Video data uploaded to the server
[1331] Step 7:
[1332] The server analyzes the received video data and integrates it with the existing highlight footage, inserting additional footage into the existing data at the appropriate time to create a more immersive final video.
[1333] Input: Uploaded user video data
[1334] Output: Final merged video data
[1335] Step 8:
[1336] The server then combines all the video data to create the final video. Specifically, it combines the edited clips, adds transitions and titles, and creates a timeline. The final video is saved in a standard format (MP4 or MOV) and uploaded to cloud storage.
[1337] Input: Final merged video data
[1338] Output: Final video file saved in cloud storage
[1339] Step 9:
[1340] The user sends a request to re-edit a specific scene to the server via their device. They enter details into a request form and send it to the server. The server then analyzes the video data again and performs a new edit. This is done using video editing software.
[1341] Input: Request for re-edit
[1342] Output: Re-edited video
[1343] Step 10:
[1344] The server saves the re-edited video back to cloud storage and provides the user with a shared link, allowing them to easily access the newly edited video.
[1345] Input: Re-edited video
[1346] Output: Video Recut files saved in cloud storage and a share link
[1347] (Application example 1)
[1348] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1349] In traditional store operations, analyzing customer behavior, providing efficient customer service, implementing security measures, and implementing effective marketing often consumes a significant amount of time and cost. It can also be difficult to grasp the situation in real time and take prompt action. This creates challenges that prevent efficient store operations and improved customer experience.
[1350] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1351] In this invention, the server includes a means for analyzing video and audio data obtained by the camera in real time to identify people's movement paths and areas where they are staying, a means for notifying detected information in real time and encouraging proactive responses, and a means for generating a heat map of customer traffic and displaying it on an interface accessible by a manager, thereby making it possible to improve the efficiency of store operations and customer satisfaction.
[1352] A "camera" is a device used to collect video and audio and transmit it to a server via wireless communication.
[1353] A "server" is a computer system that analyzes collected video and audio data and performs various processes such as identification, editing, and notification.
[1354] "Machine learning" is a technology that analyzes video and audio data to automatically recognize and identify specific patterns and highlight scenes.
[1355] "Facial recognition" is a technology for recognizing a person's face from collected video data and identifying that person.
[1356] A "highlight scene" refers to a particularly important or interesting moment or action in video or audio data.
[1357] "Wireless communication" is a technology that uses wireless LAN to send and receive data.
[1358] "User" refers to an individual or organization that shoots video data and uploads the data to a server using a terminal.
[1359] A "terminal" is a device (such as a smartphone or tablet) that a user uses to upload video data to a server.
[1360] "Traffic lines" refer to the routes that show how people move around the store.
[1361] A "stay area" refers to an area where a person or customer stays in a particular location for a certain period of time.
[1362] "Real-time notifications" are notifications sent instantly to prompt necessary action by immediately processing collected information.
[1363] "Proactive response" means predicting the situation and providing appropriate measures and services immediately.
[1364] A "heat map" is a visual display of the popularity and frequency of use of a particular area, using color coding based on collected data.
[1365] This invention is a system that uses multiple cameras to collect video and audio in real time, and a server analyzes and edits the data. This system can be applied to store operations, customer service, security measures, and marketing optimization.
[1366] Hardware and software used
[1367] Hardware:
[1368] Wireless LAN camera: Collects video and audio in real time (e.g., a typical network camera).
[1369] Server: Analyzes video and audio data and performs various processes such as identification, editing, and notification (e.g., cloud servers such as AWS EC2).
[1370] Smartphone: A device used by users to shoot and upload video data (e.g., iPhone, Android device).
[1371] Robots: Serve customers and guide them in stores (e.g., Pepper robot).
[1372] software:
[1373] Video analytics AI: Analyzes video data to identify customer movement patterns and areas of congestion (e.g., OpenCV, TensorFlow).
[1374] Speech recognition AI: Converts voice data into text and recognizes questions or requests (e.g., Google Speech-to-Text API).
[1375] Database: Stores and manages collected data (e.g., AWS RDS).
[1376] Real-time communication: Providing immediate notifications (e.g., WebSockets).
[1377] Streaming services: Real-time delivery of video data (e.g., AWS Kinesis).
[1378] System Overview
[1379] 1. Camera Initial Setup:
[1380] The camera is connected to the wireless LAN, and the terminal registers the camera's IP address, field of view, and installation location on the server.
[1381] 2. Video Collection and Transmission:
[1382] The camera collects video and audio from the entire venue in real time and transmits it to a server via wireless communication.
[1383] 3. Video and audio analysis:
[1384] The server receives the streaming data in real time and uses machine learning to analyze the video and audio data, identifying people through facial recognition and pinpointing their movement patterns and areas of concentration.
[1385] 4. Real-time notifications and actions:
[1386] The detected information is sent in real time to the store clerk's smartphone or the robot, encouraging proactive response.
[1387] 5. Heatmap and report generation:
[1388] Customer traffic heat maps are generated and displayed in the admin interface, and regular reports are generated based on the collected and analyzed data.
[1389] 6. Re-editing and final video creation:
[1390] The server re-edits specific scenes as needed and generates and serves the final video in response to the user's request.
[1391] Specific examples
[1392] Customer behavior analysis:
[1393] Cameras track customers' movements within the store and measure the time they spend in front of specific product areas, and the server then reports back to them the most popular products in the areas where customers spend the most time.
[1394] Prompt Sentence Examples
[1395] Customer behavior analysis:
[1396] - "Please identify customer movement patterns and areas where customers stay using this camera video data."
[1397] - "Please recognize the customer's question from this voice data."
[1398] This system will improve the efficiency of store operations and customer satisfaction, allowing for real-time situation assessment and rapid response.
[1399] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1400] Step 1:
[1401] Camera Defaults:
[1402] The terminal installs wireless LAN-enabled cameras in various areas of the store. The terminal registers the camera's IP address, field of view, and installation location on the server. The terminal receives the camera's physical installation location and setting information as input, and stores this information in a database as output.
[1403] Step 2:
[1404] Video collection and transmission:
[1405] The cameras collect video and audio from the entire venue in real time. This data is sent to the server via wireless communication. The system receives video and audio data from the cameras as input and sends it to the server via wireless communication. The output is treated as streaming data.
[1406] Step 3:
[1407] Video and audio analysis:
[1408] The server receives the streaming data in real time and analyzes it using video analysis AI and voice recognition AI. This allows it to identify customer movement patterns, areas where customers stay, personal identification using facial recognition, and the details of their questions and requests. It receives streaming data as input and stores the identified information (movement patterns, areas where customers stay, personal identification results, and questions) as output in a database.
[1409] Step 4:
[1410] Real-time notifications and actions:
[1411] The server notifies the detected information in real time to the store clerk's smartphone or the in-store robot, allowing the store clerk or robot to proactively respond to customers. The server receives the analysis results as input and sends notification data as output via WebSocket.
[1412] Step 5:
[1413] Generate a heatmap:
[1414] The server generates a heat map of customer traffic within the store based on data on customer movement and lingering areas. The generated heat map is displayed on an interface that can be accessed by administrators. The server receives data on movement and lingering areas as input, generates a heat map as output, and displays it on the interface.
[1415] Step 6:
[1416] Report Generation:
[1417] The server periodically generates reports based on the collected and analyzed data. The reports include customer movement patterns, areas where customers stay, questions, etc. It receives various data stored in the database as input, generates reports as output, and sends them to the administrator.
[1418] Step 7:
[1419] Re-editing and producing the final video:
[1420] The server re-edits specific scenes as needed and generates and provides the final video in response to the user's request. It receives the user's request and video data as input, generates the re-edited video as output, and provides it to the user.
[1421] In this way, real-time data collection and analysis, immediate notification and response, and visualization and reporting of analysis results will lead to more efficient store operations and improved customer satisfaction.
[1422] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1423] This invention combines a system that uses multiple cameras to collect and analyze video and audio in real time with an emotion engine. Below, we will explain the specific program processing of this system in natural language.
[1424] Camera Initial Settings
[1425] The terminal places wireless LAN-compatible cameras in various locations in the venue. The cameras are installed on the ceiling, on tabletops, and around tables, and registers information such as the camera's IP address, field of view, and installation location on the server.
[1426] Video collection and transmission
[1427] The camera begins collecting video and audio from the entire venue in real time and transmits it to a server via wireless LAN.
[1428] Video and audio analysis
[1429] The server receives the streaming data sent from the camera in real time, analyzes the video and audio, and primarily performs facial recognition to identify guests based on group photos and data acquired in advance.
[1430] Incorporating an emotion engine
[1431] The server uses an emotion engine to analyze the user's emotional state in real time based on the collected video and audio data. Based on this emotional state, highlight scenes are identified according to emotional expressions and volume changes in the video. For example, moments of cheering or crying are automatically extracted.
[1432] Video editing
[1433] The server edits the video in real time based on the identified highlight scenes, cutting out unnecessary parts and optimizing scene transitions. It also appropriately edits scenes that should be emphasized according to the emotional state identified by the emotion engine.
[1434] Attendee Video Integration
[1435] Users (attendees) upload video data taken with their smartphones or cameras to their devices. The devices then send the uploaded data to the server, which receives and analyzes it. The uploaded video data is integrated into existing highlight footage based on the analysis results of the emotion engine, creating a more moving and realistic video.
[1436] Generating the final video
[1437] The server generates the final video from all the video data, and the resulting video file is provided to the user in an easily accessible format.
[1438] Re-editing function
[1439] The user sends a request to the server via their device to re-edit a specific scene. For example, a specific request such as "I want more emphasis on the bride and groom's reaction" can be accepted. The server re-analyzes the video data based on the re-editing request and performs a new edit. The re-edited video is then provided to the user.
[1440] Specific examples
[1441] 1. Video recording of a wedding reception
[1442] The terminal places wireless LAN-compatible cameras at various locations throughout the venue and performs initial setup.
[1443] The camera captures the moment when the bride and groom appear, the guests' speeches, the toast, the cake cutting, and the expressions of the many guests in real time, and sends the video and audio data to the server.
[1444] The server analyzes the transmitted data and identifies the bride and groom and important guests using facial recognition, analyzes their emotional state using an emotion engine, and identifies highlight scenes based on emotional expressions and volume changes.
[1445] The server edits the video in real time based on the identified highlights. For example, the emotion engine detects the bride's tears of joy and emphasizes these scenes.
[1446] Users (attendees) upload videos they have taken with their smartphones, which are then received and analyzed by the server. The uploaded video is then integrated with existing edited footage based on the analysis results of the emotion engine, allowing for a richer reproduction of the bride and groom's reactions and the guests' emotional moments.
[1447] The server generates the final edited video and makes it ready to be served.
[1448] In this way, a system that combines an emotion engine enables high-quality video recording and editing that reflects the user's emotional state.
[1449] The processing flow will be explained below.
[1450] Step 1:
[1451] The terminal installs wireless LAN-compatible cameras in advance at various locations in the venue and registers information such as the camera's IP address, field of view, and installation location on the server.
[1452] Step 2:
[1453] The cameras begin collecting video and audio from the entire venue in real time and streaming it to a server via Wi-Fi.
[1454] Step 3:
[1455] The server receives the streaming data sent from the camera in real time and analyzes the video and audio.
[1456] Step 4:
[1457] The server performs facial recognition based on group photos and data acquired in advance to identify guests.
[1458] Step 5:
[1459] The server uses AI to analyze video and audio data, detecting emotional expressions and changes in volume to identify highlight scenes.
[1460] Step 6:
[1461] The server uses an emotion engine to analyze the user's emotional state in real time, and identifies important scenes in the video based on the emotional state.
[1462] Step 7:
[1463] The server then prioritizes editing the identified highlights and optimizes scene transitions by cutting out unnecessary parts, for example, by emphasizing touching moments between the bride and groom or scenes that elicit cheers of joy.
[1464] Step 8:
[1465] Users (attendees) upload video data taken with their smartphones or cameras to the device.
[1466] Step 9:
[1467] The device sends the uploaded video data of attendees to the server, which receives and analyzes it.
[1468] Step 10:
[1469] The server then integrates the uploaded attendee video data with existing highlight footage based on the analysis results of the emotion engine, creating a more moving and realistic video.
[1470] Step 11:
[1471] The server generates the final video from all the video data and provides the video file in an easily accessible format for the user.
[1472] Step 12:
[1473] The user sends a request to the server via their device to re-edit a specific scene, such as "I want more emphasis on the reaction of the bride and groom."
[1474] Step 13:
[1475] The server re-analyzes the video data based on the re-editing request and performs a new edit. The re-edited video is then provided to the user.
[1476] Example 2
[1477] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1478] Modern events require efficient recording and editing of many important moments. However, doing so manually is a time-consuming and labor-intensive task. Furthermore, integrating footage shot separately by attendees is tedious and does not necessarily produce natural-looking results. Furthermore, capturing emotional expressions and volume changes in the footage and identifying moving moments and important scenes is a difficult challenge. Therefore, there is a need for a system that can automatically collect, analyze, and edit footage in real time, and efficiently integrate additional footage from users.
[1479] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1480] In this invention, the server includes means for collecting video and audio from a camera in real time, means for transmitting the collected video and audio to the server via wireless communication, means for the server to perform individual identification using facial recognition technology, means for analyzing the video and audio using machine learning to identify important scenes, means for editing in real time based on the important scenes identified by the server, means for uploading video data shot by a user to the server using a terminal, means for integrating the uploaded user's video data with existing important scene videos, means for generating a final video using all the video data, means for re-editing based on a request to re-edit a specific scene, and means for outputting the generated final video and re-edited video. This makes it possible to automatically and efficiently collect, analyze, and edit video, and naturally integrate additional user video.
[1481] A "camera" is a device for collecting video and audio and capable of transmitting data to a server via a wireless communication network.
[1482] "Server" is a central processing unit for receiving, analyzing, and editing collected video and audio data.
[1483] "Facial recognition technology" refers to algorithms and techniques for identifying people from video data.
[1484] "Machine learning" is a technology that analyzes large amounts of data to learn patterns and characteristics and make predictions and distinctions.
[1485] An "important scene" refers to a particularly noteworthy or moving moment in a video, and is identified by emotional expressions or changes in volume.
[1486] "Real-time editing" is the process of editing scenes that are deemed important at the moment as video data is collected.
[1487] "User-recorded video data" refers to video files that event attendees record themselves using their smartphones or cameras and later upload to the system.
[1488] "Uploading" refers to the act of sending video data taken by a user to a server.
[1489] "Merge" is the process of combining uploaded footage with existing footage and key scenes.
[1490] The "final video" is a video file in the final format after all video data has been edited and compiled.
[1491] "Re-editing" is a process in which already edited video data is re-edited based on a specific user request.
[1492] "Generated final video or re-edited video" includes final version video files automatically generated by the system and re-edited video files based on user requests.
[1493] A "wireless communication enabled camera" is a camera that has the ability to transmit data using a wireless communication network.
[1494] A "communications network" is an infrastructure for transmitting and receiving data between multiple devices using wireless LAN or other wireless communication protocols.
[1495] "Emotional expression" refers to the psychological state and emotional expression obtained by analyzing facial expressions in the video.
[1496] "Volume change" refers to an increase or decrease in volume detected by analyzing the audio data in the video.
[1497] The present invention combines a system that utilizes multiple wireless communication enabled cameras to collect and analyze video and audio in real time with an emotion engine. Specific embodiments for implementing this system are described below.
[1498] First, wireless communication-enabled cameras are installed in the venue. The device (e.g., laptop or tablet PC) sets the camera's IP address, field of view, and installation location, and registers this data on the server. The cameras are installed on the ceiling, on tabletops, around tables, etc., and are positioned so that they cover the entire venue. At this point, the cameras and server are connected using a wireless communication network (such as wireless LAN).
[1499] Next, the camera starts collecting video and audio in real time. The camera also collects audio data using its built-in microphone and transmits this data wirelessly to the server. The server uses a high-performance CPU and GPU to receive and analyze the streaming data sent from the camera in real time.
[1500] The server first uses facial recognition technology (e.g., OpenCV or dlib) to identify individuals from the collected video data, then uses an algorithm to analyze the audio data, detecting specific keywords and volume changes. The identified data is then stored in a database.
[1501] Furthermore, the server uses an emotion engine (e.g., Emotion API or Affectiva) to analyze the user's emotional state from video and audio data. The emotion engine analyzes facial expressions and audio tones for each frame to identify emotional states such as smiling or crying faces. This allows highlight scenes (important scenes) to be automatically identified.
[1502] Based on the identified highlight scenes, the server edits the video in real time. For editing, editing software (e.g., FFmpeg) is used to cut out unnecessary parts and ensure smooth scene transitions. Scenes that should be emphasized are appropriately edited according to emotional expressions and volume changes.
[1503] In addition, users (attendees) can upload video data taken with their smartphones or cameras to their devices, which then send the uploaded data to the server. The server receives the data, analyzes it using an emotion engine, and integrates it into existing highlight footage. This creates a more moving and realistic video.
[1504] Finally, the server generates the final video from all the video data, and the completed video file is uploaded to cloud storage and provided to users in a format that is easy to access (e.g., MP4, AVI).
[1505] Furthermore, users can send requests to the server via their devices to re-edit specific scenes. For example, a specific request could be made to "emphasize the bride and groom's reactions more." The server then re-analyzes the original editing data based on the re-editing request and performs a new edit. The re-edited video is also uploaded to cloud storage and provided to the user.
[1506] Specific examples
[1507] 1. Video recording of a wedding reception
[1508] The terminals are used to place wireless communication-enabled cameras at various locations in the venue, and the initial settings are performed. The camera's IP address, field of view, and installation location are registered on the server.
[1509] The camera captures the moment when the bride and groom appear, the guests' speeches, the toast, the cake cutting, and the expressions of the many guests in real time, and sends the video and audio data to the server.
[1510] The server analyzes the transmitted data and uses facial recognition to identify the bride and groom and important guests, and an emotion engine to analyze their emotional state and identify highlight scenes based on emotional expressions and volume changes.
[1511] The server edits the video in real time based on the identified highlights. For example, the emotion engine detects the bride's tears of joy and emphasizes these scenes.
[1512] Users (attendees) upload videos they have taken with their smartphones, which are then received and analyzed by the server. The uploaded video is then integrated with existing edited footage based on the analysis results of the emotion engine, allowing for a richer reproduction of the bride and groom's reactions and the guests' emotional moments.
[1513] The server generates the final edited video and makes it ready to be served. The video file is uploaded to cloud storage in MP4 format and a download link is provided to the user.
[1514] Example prompts to input to the generative AI model
[1515] "I want to create a video that selects and edits the most touching scenes from a wedding reception."
[1516] "Create a highlight video that highlights the bride and groom's reactions and the guests' joyful moments."
[1517] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1518] Step 1: Initial camera setup
[1519] The terminal places wireless communication-enabled cameras in the venue. Specifically, it sets the camera's IP address, field of view, and installation location information, and registers this data on the server. For example, the terminal checks the network connection of each camera from a settings screen and enters the setting information. The input is the camera's physical installation location, field of view, and IP address. The output is the setting information stored in the server's database.
[1520] Step 2: Collect and transmit video and audio
[1521] The camera starts collecting video and audio in real time and sends them to the server via wireless communication. Specifically, the camera captures frames every second and collects audio data. The input is the video and audio acquired from the camera's sensor. The output is data sent to the server in streaming format.
[1522] Step 3: Video and audio analysis
[1523] The server receives and analyzes streaming data sent from the camera in real time. Individual identification is performed using facial recognition technology (e.g., OpenCV or dlib). The input is video data sent from the camera. Specifically, the captured frames are run through a facial recognition algorithm to detect the position of the face and its features. The output is information about the identified person and feature data. Audio data is also analyzed to detect specific keywords and volume changes. The input is audio data sent from the camera. The output is keyword detection information and volume change information as the analysis results.
[1524] Step 4: Incorporating the Emotion Engine
[1525] The server uses an emotion engine (e.g., Emotion API or Affectiva) to analyze the user's emotional state from video and audio data. The input is the video and audio feature data analyzed in the previous step. Specifically, the emotion engine analyzes facial expressions and audio tone for each frame to identify emotional states such as laughter or tears. The output is data related to the emotional state.
[1526] Step 5: Edit your footage
[1527] The server edits the video in real time based on the highlight scenes identified by the emotion engine. Specifically, it uses editing software (e.g., FFmpeg) to cut out unnecessary parts and smooth scene transitions. The input is the identified highlight scene information and the original video data. The output is the edited video data.
[1528] Step 6: Integrate attendee videos
[1529] Users (attendees) upload video data taken with their smartphones or cameras to their terminals. The terminals then send the uploaded data to the server. Specifically, the terminals manage the process of selecting video files and transferring them to the server. The input is the video data uploaded from the user's device. The output is additional video data stored on the server. The server analyzes the data using an emotion engine and integrates it into the existing highlight video. The input is the newly uploaded video data and the existing highlight video data. The output is the integrated video data.
[1530] Step 7: Generate the final video
[1531] The server generates the final video from all the video data. Specifically, editing software is used to combine all the scenes and create the final video file. The input is the merged video data. The output is the final video file (e.g., MP4 format). The generated file is uploaded to cloud storage so that users can access it.
[1532] Step 8: Re-edit function
[1533] The user sends a request to the server via their device to re-edit a specific scene. For example, a request may be made to "emphasize the bride and groom's reaction more." The server re-analyzes the original video data based on the re-editing request and performs a new edit. Specifically, the server re-processes the original video data according to the re-editing instructions and edits it to emphasize the necessary parts. The input is the re-editing request and the original video data. The output is a re-edited video file, which is also uploaded to cloud storage and a download link is provided to the user.
[1534] (Application example 2)
[1535] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1536] There is a need for a system that can analyze customer emotions in real time and provide appropriate feedback to store staff to provide more effective customer service and improve customer satisfaction. However, existing systems lack the means to analyze customer emotions and behavior in real time, making it difficult for staff to respond immediately.
[1537] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting video and audio from a camera in real time, means for transmitting the collected video and audio to the server via wireless communication, means for the server to identify people using biometric authentication, means for the server to analyze the video and audio using a generative AI model and identify key scenes, means for editing in real time based on the key scenes identified by the server, means for uploading video data taken by participants to the server using a terminal, means for integrating the uploaded video data of participants with existing key scene videos, means for generating a final video using all video data, means for re-editing based on a request to re-edit a specific scene, means for outputting the generated final video and re-edited video, and means for the smart device to analyze customer emotions in real time and provide feedback to staff. This makes it possible to analyze customer emotions in real time within the store and provide prompt and appropriate support.
[1538] A "camera" is a device that collects video and audio and converts it into electrical signals.
[1539] "Wireless communication" is a means of sending and receiving data without using cables, and is a technology that uses radio signals.
[1540] A "server" is a computer system that processes and stores data on a network and provides various services.
[1541] "Biometric authentication" is a technology that identifies individuals using biological characteristics such as their face, fingerprints, and irises.
[1542] A "generative AI model" is an algorithm or program that uses artificial intelligence to analyze data and generate new information or results.
[1543] "Video" is visual data generated by optical means and collected by a camera or other imaging device.
[1544] "Sound" refers to a sound signal transmitted by vibrations in the air, and is collected by a sound collection device such as a microphone.
[1545] An "important scene" refers to a scene or moment in video or audio data that is particularly noteworthy.
[1546] "Real-time" refers to data processing and analysis without delay, providing results almost instantly.
[1547] A "terminal" is a device that connects to a network and transmits, receives, and processes data.
[1548] "Integration" refers to combining multiple data or elements into one.
[1549] A "participant" is an individual or group who is actively involved in an event, activity, or project.
[1550] A "smart device" is an advanced electronic device that has communication capabilities and processes and presents information.
[1551] This invention describes a system that analyzes customer sentiment in brick-and-mortar stores and provides real-time feedback to store staff. The system is composed of a combination of cameras, audio collection devices, wireless communication connections, a server, smart devices, and a generative AI model.
[1552] Hardware and Software Configuration
[1553] Camera and audio collection device
[1554] Cameras are installed in various locations throughout the store and collect video and audio recordings of customers in real time. They have wireless communication capabilities and transmit video and audio data to a server. Specifically, store staff wear smart glasses (e.g., Google Glass or Vuzix M400) to collect customer interactions.
[1555] server
[1556] The server receives the collected video and audio data via wireless communication and analyzes it. This analysis involves identifying customers using biometric authentication technology and performing sentiment analysis using generative AI models (e.g., Affectiva, Microsoft Azure Cognitive Services). Furthermore, key scenes are identified based on the video and audio data, and editing is performed in real time.
[1557] Smart Devices
[1558] The smart device (smart glasses) displays the analysis results from the server, allowing store staff to instantly understand the customer's emotional state and respond appropriately.
[1559] Data processing and calculation
[1560] The server does the following:
[1561] 1. Receive video and audio data and identify customers using biometric technology.
[1562] 2. Use generative AI models to analyze customer sentiment in real time from video and audio data.
[1563] 3. Identify key scenes based on the customer's emotional state, edit in real time, and provide feedback to smart devices.
[1564] Specific examples
[1565] In practice, the following scenarios can be considered when using this system:
[1566] Situation 1: Dealing with customers in the fitting room
[1567] Example prompt sentence:
[1568] "If a staff member wearing glasses carefully observes a customer entering a fitting room and detects that the customer has a troubled expression."
[1569] What actually happens:
[1570] Smart glasses observe customers in fitting rooms.
[1571] The emotion engine analyzes the customer's facial expression to determine whether they are confused.
[1572] "The customer is in trouble" is displayed in real time on the smart glasses display.
[1573] Staff immediately went to the fitting room to provide the customer with appropriate assistance.
[1574] Situation 2: Serving customers at the counter
[1575] Example prompt sentence:
[1576] "When a staff member wearing glasses observes a customer who comes to a product counter, they detect that the customer has a satisfied expression."
[1577] What actually happens:
[1578] Smart glasses observe customers at the counter.
[1579] The emotion engine analyzes the customer's facial expression to determine whether they are satisfied.
[1580] "Customer is satisfied" is displayed in real time on the smart glasses display.
[1581] Staff will further enhance customer service, suggesting additional products and offering special services to customers.
[1582] This makes it possible to analyze customer sentiment in real time and provide appropriate support.
[1583] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1584] Step 1:
[1585] The cameras and audio collection devices begin to collect video and audio data in real time. This data captures how customers move around the store and what they say. The input is the video and audio captured by the cameras and microphones, and the output is the data transmitted wirelessly to the server.
[1586] Step 2:
[1587] The server receives and analyzes the data sent from the camera and audio collection device. Specifically, it uses a biometric authentication algorithm to identify customers from the video data. The input is the video and audio data sent from the camera and audio collection device, and the output is the ID information of the identified customer.
[1588] Step 3:
[1589] The server uses a generative AI model to analyze the customer's emotions from the received video and audio data, including facial expressions and tone of voice. The input is the identified customer's video and audio data, and the output is the customer's emotional state (e.g., interest, confusion, satisfaction, etc.).
[1590] Step 4:
[1591] The server identifies key moments in real time based on the customer's emotional state. Here, a generative AI model detects emotional peaks and changes and extracts specific events (e.g., excitement, confusion). The input is the customer's emotional state data, and the output is the timestamp and context of the identified key moments.
[1592] Step 5:
[1593] The server edits the video in real time based on the identified important scenes, cutting out unnecessary parts and emphasizing the important scenes. The input is the timestamp and context of the identified important scenes, and the output is the edited video clip.
[1594] Step 6:
[1595] The server sends the edited video clip to the smart device and provides feedback to the store staff in real time. The input is the edited video clip and the emotion analysis results, and the output is the feedback information displayed on the smart device's display.
[1596] Step 7:
[1597] Store staff wearing smart devices respond based on the feedback provided by the server. Specifically, if a customer's confusion is detected, the staff responds immediately and provides appropriate support. The input is the feedback information displayed on the smart device, and the output is the staff's immediate response.
[1598] Step 8:
[1599] The server stores all collected data and later analyzes it for long-term patterns and trends of customer behavior. The input is the entire video and audio data collected daily, and the output is regular reports and analysis results, providing information to support strategic decision-making.
[1600] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1601] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1602] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1603] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1604] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1605] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1606] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1607] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1608] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1609] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1610] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1611] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1612] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1613] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1614] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1615] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1616] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1617] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1618] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1619] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1620] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1621] The following is further disclosed regarding the above embodiment.
[1622] (Claim 1)
[1623] A means of collecting video and audio from the camera in real time;
[1624] A means to transmit the collected video and audio to a server via wireless LAN,
[1625] A means for the server to identify guests using face recognition;
[1626] The server uses AI to analyze the video and audio and identify highlight scenes.
[1627] A means for editing in real time based on the highlight scenes identified by the server;
[1628] A means for uploading video data taken by attendees to a server using a terminal;
[1629] a means for integrating uploaded attendee video data into existing highlight footage;
[1630] a means for generating a final video using all of the video data;
[1631] a means for performing re-editing based on a request to re-edit a particular scene;
[1632] The system includes a means for outputting the generated final video or re-edited video.
[1633] (Claim 2)
[1634] 2. The system according to claim 1, wherein a plurality of network cameras are connected to a server via a wireless LAN.
[1635] (Claim 3)
[1636] The system according to claim 1, characterized in that the AI identifies highlight scenes based on emotional expressions and volume changes.
[1637] "Example 1"
[1638] (Claim 1)
[1639] A means of collecting video and audio from the camera in real time;
[1640] a means for transmitting the collected video and audio to a server via wireless communication;
[1641] A means for the server to identify guests using face recognition;
[1642] A means for the server to use artificial intelligence to analyze the video and audio and identify highlight scenes;
[1643] A means for editing in real time based on the highlight scenes identified by the server;
[1644] A means for uploading video data taken by attendees to a server using a terminal;
[1645] a means for integrating uploaded attendee video data into existing highlight footage;
[1646] a means for generating a final video using all of the video data;
[1647] a means for performing re-editing based on a request to re-edit a particular scene;
[1648] The system includes a means for outputting the generated final video or re-edited video.
[1649] (Claim 2)
[1650] 2. The system according to claim 1, wherein a plurality of network cameras are connected to a server via wireless communication.
[1651] (Claim 3)
[1652] The system according to claim 1, characterized in that the artificial intelligence identifies highlight scenes based on emotional expressions and volume changes.
[1653] "Application Example 1"
[1654] (Claim 1)
[1655] A means of collecting video and audio from the camera in real time;
[1656] a means for transmitting the collected video and audio to a server via wireless communication;
[1657] A means for the server to identify a person using face recognition;
[1658] A means for the server to use machine learning to analyze video and audio and identify highlight scenes;
[1659] A means for editing in real time based on the highlight scenes identified by the server;
[1660] A means for uploading video data taken by a user to a server using a terminal;
[1661] a means for integrating uploaded user video data into existing highlight videos;
[1662] a means for generating a final video using all of the video data;
[1663] a means for performing re-editing based on a request to re-edit a particular scene;
[1664] A means for outputting the generated final video or re-edited video;
[1665] A method for analyzing video and audio data obtained by cameras in real time to identify people's movements and areas where they stay.
[1666] Real-time notification of detected information and a means to encourage proactive response.
[1667] A means to generate and display heat maps of customer traffic in an administrator-accessible interface;
[1668] means for periodically generating reports based on the collected and analyzed data;
[1669] A means to propose optimization of store layout and product placement based on collected data,
[1670] A system that includes a means of recognizing customer questions and requests from voice data and prompting a response.
[1671] (Claim 2)
[1672] 2. The system according to claim 1, wherein a plurality of network cameras are connected to a server via wireless communication.
[1673] (Claim 3)
[1674] The system of claim 1, characterized in that machine learning identifies highlight scenes based on emotional expressions and volume changes.
[1675] "Example 2: Combining Emotion Engines"
[1676] (Claim 1)
[1677] A means of collecting video and audio from the camera in real time;
[1678] a means for transmitting the collected video and audio to a server via wireless communication;
[1679] A means for the server to identify individuals using face recognition technology;
[1680] The server uses machine learning to analyze video and audio and identify important scenes.
[1681] A means to edit in real time based on important scenes identified by the server,
[1682] A means for uploading video data taken by a user to a server using a terminal;
[1683] A means for integrating uploaded user video data into existing key scene video;
[1684] a means for generating a final video using all of the video data;
[1685] a means for performing re-editing based on a request to re-edit a particular scene;
[1686] The system includes a means for outputting the generated final video or re-edited video.
[1687] (Claim 2)
[1688] 2. The system according to claim 1, wherein a plurality of wireless communication-enabled cameras are connected to a server via a communication network.
[1689] (Claim 3)
[1690] The system of claim 1, characterized in that the machine learning model identifies important scenes based on emotional expressions and volume changes.
[1691] "Application example 2 when combining emotion engines"
[1692] (Claim 1)
[1693] A means of collecting video and audio from the camera in real time;
[1694] a means for transmitting the collected video and audio to a server via wireless communication;
[1695] A means for the server to identify a person using biometric authentication;
[1696] The server uses a generative AI model to analyze video and audio and identify key scenes.
[1697] A method for editing in real time based on important scenes identified by the server,
[1698] A means for participants to upload the video data they have taken to a server using their terminals;
[1699] A means of integrating uploaded participant video data with existing key scene video;
[1700] a means for generating a final video using all of the video data;
[1701] a means for performing re-editing based on a request to re-edit a particular scene;
[1702] A means for outputting the generated final video or re-edited video;
[1703] A system that includes a means for smart devices to analyze customer sentiment in real time and provide feedback to staff.
[1704] (Claim 2)
[1705] 2. The system according to claim 1, wherein a plurality of network cameras are connected to a server via wireless communication.
[1706] (Claim 3)
[1707] The system of claim 1, characterized in that the generative AI model identifies important scenes based on emotional expressions and volume changes. [Explanation of symbols]
[1708] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of collecting video and audio from the camera in real time; A means to transmit the collected video and audio to a server via wireless LAN, A means for the server to identify guests using face recognition; The server uses AI to analyze the video and audio and identify highlight scenes. A means for editing in real time based on the highlight scenes identified by the server; A means for uploading video data taken by attendees to a server using a terminal; a means for integrating uploaded attendee video data into existing highlight footage; a means for generating a final video using all of the video data; a means for performing re-editing based on a request to re-edit a particular scene; The system includes a means for outputting the generated final video or re-edited video.
2. 2. The system according to claim 1, wherein a plurality of network cameras are connected to a server via a wireless LAN.
3. The system according to claim 1, characterized in that the AI identifies highlight scenes based on emotional expressions and volume changes.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A