System

By extracting and compressing feature data from real-time video and audio in online conferences, the system addresses network strain and environmental impact, achieving efficient and cost-effective data transmission.

JP2026018007APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024119068
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Modern online conference services face challenges with strained network resources and increased transmission costs due to the transmission and reception of large amounts of real-time video and audio data, which also have a significant environmental impact.

Method used

A system that extracts feature data from real-time video and audio, compresses it, and transmits it as packets, using a receiving terminal with generative AI to recreate the video and audio, reducing the amount of data sent and received.

Benefits of technology

This approach significantly reduces network resource loads, transmission costs, and environmental impact while providing high-quality real-time video and audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026018007000001_ABST
    Figure 2026018007000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The method of claim 1, further comprising the step of generating real-time video and audio using a generation AI based on the feature information, wherein the step of generating real-time video and audio includes the step of generating real-time video and audio using the generation function.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Modern online conference services require the transmission and reception of large amounts of data to deliver real-time video and audio, resulting in problems such as strained network resources and increased transmission costs. Furthermore, transmitting and receiving large amounts of data consumes a lot of energy, placing a significant burden on the environment. There is a need to solve these issues and provide an efficient, low-cost, and environmentally friendly video streaming system. [Means for solving the problem]

[0005] The present invention provides a system that includes a transmitting terminal that extracts feature data from real-time video and audio data, compresses the feature data, and transmits the transmission packets to a receiving terminal. The receiving terminal also includes a generating AI that generates real-time video and audio based on the feature data. This configuration significantly reduces the amount of data sent and received, reducing network resource loads, transmission costs, and environmental impact.

[0006] A "transmitting terminal" is a device that captures video and audio, extracts feature data, and transmits it to a receiving party in an online conference or the like.

[0007] "Real-time video and audio data" refers to video and audio data that is transmitted instantaneously and records ongoing conversations and actions.

[0008] "Feature data" refers to the minimum information extracted from video and audio data, and specifically includes facial landmark points and audio frequency spectrum characteristics.

[0009] "Compression" is a method of efficiently reducing data size in order to reduce the original data volume.

[0010] A "transmission packet" is a unit of data transmitted over a network, and is a packet made up of compressed data containing feature data.

[0011] The "recipient's terminal" is a device that receives data sent from the sender's terminal and generates real-time video and audio using generation AI.

[0012] "Generative AI" refers to algorithms and models that use artificial intelligence technology to generate real-time video and audio from feature data.

[0013] An "algorithm" is a set of procedures or computational rules for solving a particular problem.

[0014] A "machine learning model" is a mathematical or statistical model used to learn and perform a specific task using large amounts of data. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] The system of the present invention utilizes generative AI to generate video and audio in real time based on feature data in order to improve the efficiency of data transmission and reception in online conferences and video streaming.

[0037] Program processing overview

[0038] User connection and pre-transmission of data

[0039] The server receives profile images and voice samples of users when they connect to an online conference. During this process, users submit their own images and voice data introducing themselves, which are then stored on the server.

[0040] Extracting and sending feature data

[0041] The transmitting device captures the user's camera video and microphone audio in real time and extracts feature data from them, including facial landmarks (such as the positions of the eyes, nose, and mouth) and frequency spectrum characteristics of the voice.

[0042] The terminal compresses the extracted feature data and transmits it to the server as an efficient transmission packet.

[0043] Video Creation and Distribution

[0044] The server analyzes the feature data received from the sending device and transmits it to the receiving device in an appropriate format, allowing the generating AI to recreate the video and audio in real time.

[0045] The receiving device runs the generative AI using the feature data sent from the server to generate video and audio in real time, which is then displayed to the receiving user, providing an experience similar to that of a regular online meeting.

[0046] Specific examples

[0047] 1. Connection and Pre-Data Transmission

[0048] User A connects to the online conference, and the server receives the profile picture and self-introduction voice of User A. User A's terminal transmits this data to the server.

[0049] 2. Extracting and sending feature data

[0050] The transmitting terminal (User A's device) captures User A's video and audio in real time during the conference and extracts feature data such as facial position information and audio frequency characteristics. This feature data is efficiently compressed and sent to the server as a transmission packet.

[0051] 3. Video Creation and Distribution

[0052] The server analyzes the received feature data and sends it in an appropriate format to the receiving device (User B's device). The receiving device inputs the transmitted feature data into the generation AI, which generates real-time video and audio of User A. User B then watches the generated video and audio, enjoying a natural online meeting experience.

[0053] In this way, the system of the present invention can provide real-time video and audio while significantly reducing the amount of data transmitted and received and easing the strain on network resources, resulting in cost reduction and a reduction in the environmental impact.

[0054] The processing flow will be explained below.

[0055] Step 1:

[0056] A user connects to an online meeting.

[0057] The user launches the online meeting application and connects to the server.

[0058] Step 2:

[0059] The server receives the advance data.

[0060] The server receives a profile image and voice sample from the user, which includes a photo of the user's face and an audio introduction.

[0061] Step 3:

[0062] The user submits the advance data.

[0063] Users send their profile picture and voice sample from their device to the server, and this data is sent when the meeting is initially connected.

[0064] Step 4:

[0065] The device captures video and audio.

[0066] The sending device captures the user's camera video and microphone audio in real time, allowing conversations and actions to be recorded.

[0067] Step 5:

[0068] The device extracts the feature data.

[0069] The device extracts facial landmarks (such as the positions of the eyes, nose, and mouth) and gestures from the captured video, and also extracts frequency spectrum characteristics from the audio data.

[0070] Step 6:

[0071] The terminal compresses the feature data.

[0072] The terminal efficiently compresses the extracted feature data to generate transmission packets, which significantly reduces the amount of data.

[0073] Step 7:

[0074] The terminal transmits the feature data to the server.

[0075] The terminal transmits a transmission packet to the server in real time, and the transmission packet includes compressed feature data.

[0076] Step 8:

[0077] The server analyzes the feature data.

[0078] The server parses the received feature data and decodes it into the appropriate format, making the data usable by the receiving device.

[0079] Step 9:

[0080] The server transmits the data to the receiving terminal.

[0081] The server transmits the analyzed feature data to the receiving terminal, which then generates the data in real time.

[0082] Step 10:

[0083] The receiving device launches the generation AI.

[0084] The receiving device activates the generation AI based on the transmitted feature data and inputs the data, which starts the generation process.

[0085] Step 11:

[0086] Generative AI generates real-time video and audio.

[0087] The generative AI uses the feature data to generate real-time video and audio that replicates the movements and mouth movements of the sending user.

[0088] Step 12:

[0089] The receiving terminal displays the generated video and audio to the user.

[0090] The receiving terminal displays the generated video and audio to the user in real time, and the user can see and hear the video and audio of other participants just like in a regular online conference.

[0091] This system significantly reduces the amount of data sent and received, easing the load on network resources and resulting in cost savings and a reduced environmental impact.

[0092] Example 1

[0093] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0094] In online conferences and video streaming, sending and receiving real-time video and audio data requires a large amount of network resources, resulting in communication delays and quality degradation. Furthermore, existing technologies lack efficient methods for generating high-quality video and audio in real time. This results in poor user experience and increases costs and environmental impact.

[0095] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0096] In this invention, the server includes a means for a user to connect to an online conference and transmit a profile image and a voice sample, a transmitting terminal including means for extracting feature data including facial landmarks and voice frequency spectrum characteristics from real-time video and audio data, a means for compressing the feature data into a transmission packet, a means for transmitting the transmission packet to the server and for the server to convert the feature data into an appropriate format, a means for the server to transmit the converted data to a receiving terminal, and a means for the receiving terminal to generate real-time video and audio using a generative AI model based on the feature data. This significantly reduces the amount of data transmitted and received, enabling real-time generation of high-quality video and audio while minimizing communication delays.

[0097] "Users" are typical participants who use online meetings and video streaming services.

[0098] An "online conference" refers to a conference or meeting that takes place over the Internet in real time, sharing audio and video.

[0099] A "profile image" is a still image that a user uses to indicate their personal information or identity.

[0100] A "voice sample" is a short piece of voice data that a user provides to the server, used for self-introduction and identification.

[0101] The "transmitting terminal" is a device used by a user (for example, a smartphone or a personal computer) that captures video and audio using a camera and microphone.

[0102] "Feature data" are specific digital indices or characteristics (e.g., facial landmarks or frequency spectrum characteristics of voice) extracted from captured video and audio data.

[0103] "Landmarks" are characteristic points that indicate the position information of each part of the face (eyes, nose, mouth, etc.).

[0104] The "frequency spectrum characteristics" are data that indicate the strength of each frequency component in an audio signal.

[0105] A "server" is a central system that manages and processes data sent and received from sending and receiving terminals.

[0106] A "transmission packet" is a small unit of compressed data containing feature data that is transmitted over a network.

[0107] The "receiving terminal" is a device that receives the feature data sent from the server and reproduces real-time video and audio using a generation AI.

[0108] A "generative AI model" is an artificial intelligence technology that uses machine learning algorithms to generate video and audio from feature data.

[0109] "Compression" is a technique for reducing data size and is used to improve transmission efficiency.

[0110] "Real-time" refers to instantaneous processing and communication with little or no delay.

[0111] MODE FOR CARRYING OUT THE INVENTION

[0112] The system of the present invention utilizes a generative AI model to generate video and audio in real time based on feature data in order to improve the efficiency of data transmission and reception in online meetings and video streaming. A specific embodiment of this system is described below.

[0113] Hardware and software used

[0114] Hardware: A computer device with a compatible camera, microphone, and internet connection

[0115] Software: Generative AI models (e.g., GPT-3, DALL-E), libraries for video and audio capture (e.g., OpenCV, PyAudio)

[0116] A natural language description of the program

[0117] User connection and pre-transmission of data

[0118] A user opens an online conference app on their computer and connects to the server. The user then sends a profile picture and a self-introduction voice sample. During this process, the user follows the application's instructions to select a profile picture they have taken in advance, record a self-introduction voice using a microphone, and send it to the server. The server then stores the received data.

[0119] Real-time extraction and transmission of feature data

[0120] The sending device (e.g., User A's device) uses a camera and microphone to capture User A's video and audio in real time. Specifically, it uses a computer vision library such as OpenCV to detect facial landmarks and an audio analysis library such as PyAudio to analyze the frequency spectrum characteristics of the audio. These feature data are efficiently compressed and sent to the server as transmission packets.

[0121] Feature data analysis and resubmission

[0122] The server analyzes the feature data received from the sending device and converts the data into the appropriate format. This involves using generative AI models to calculate additional data needed to generate the video and audio. For example, it generates data to complete facial details and audio. The server then transmits the converted data to the receiving device.

[0123] Real-time video and audio generation

[0124] The receiving device (e.g., User B's device) receives the feature data and complementary data sent from the server, and the internal generative AI model is activated. For example, DALL-E generates a video reproducing User A's face based on the facial landmark data and complementary data, and simultaneously generates real-time audio based on the audio sample. The generated video and audio are output to the screen and speakers, just like a regular video conferencing app.

[0125] Specific examples

[0126] 1. Connection and Pre-Data Transmission

[0127] User A connects to the online conference, and the server receives the profile picture and self-introduction voice of User A. User A's terminal transmits this data to the server.

[0128] 2. Real-time extraction and transmission of feature data

[0129] The transmitting terminal (User A's device) captures User A's video and audio in real time during the conference and extracts feature data such as facial position information and audio frequency characteristics. This feature data is efficiently compressed and sent to the server as a transmission packet.

[0130] 3. Analysis and retransmission of feature data

[0131] The server analyzes the received feature data and sends it in an appropriate format to the receiving device of User B. User B's device inputs the transmitted feature data into the generation AI, which generates real-time video and audio of User A.

[0132] Prompt Sentence Examples

[0133] Data sent from User A's device to the server: "Send profile picture and voice sample"

[0134] The process in which the server sends feature data to User B's device: "Generate video and audio based on facial landmarks and audio spectrum characteristics."

[0135] This system significantly reduces the amount of data transmitted and received, alleviating pressure on network resources while providing high-quality real-time video and audio, resulting in cost savings and a lighter environmental impact.

[0136] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0137] Step 1:

[0138] User connection and pre-transmission of data

[0139] A user opens an online conference app on their computer and connects to the server. The user provides a profile image and a self-introduction voice sample as input, which is then sent to the server. The input image and voice data are stored on the server. For example, the user may upload a face photo taken with a camera or a voice recording with a microphone. This process is performed to collect initial data that will serve as the basis for subsequent feature data generation.

[0140] Step 2:

[0141] Real-time extraction and transmission of feature data

[0142] The sending terminal (user's device) uses a camera and microphone to capture the user's video and audio in real time. This becomes the input data. Specifically, a computer vision library such as OpenCV is used to detect facial landmarks (the positions of the eyes, nose, mouth, etc.), and an audio analysis library such as PyAudio is used to analyze the frequency spectrum characteristics of the audio. This feature data is compressed and sent to the server as a transmission packet. For example, this includes operations to obtain the position information of each part of the face from the camera video and calculate the strength of specific frequency components from the audio signal. The output is compressed feature data.

[0143] Step 3:

[0144] Feature data analysis and resubmission

[0145] The server receives feature data from the sending device as input. It analyzes this data and uses a generative AI model (e.g., GPT-3 or DALL-E) to calculate the necessary completion data for video and audio generation. Specifically, it generates completion data for facial details and audio. This data is then compressed again and sent to the receiving device in an appropriate format. For example, this may involve the generative AI completing fine details of the user's face to generate high-quality video data. The output is a packet containing the completion data and the original feature data.

[0146] Step 4:

[0147] Real-time video and audio generation

[0148] The receiving device (User B's device) receives the feature data and complementary data from the server as input. An internal generative AI model is activated and generates video and audio in real time based on this data. Specifically, DALL-E generates video recreating User A's face based on facial landmark data, and generates real-time audio based on audio samples. The output is the generated video and audio, which are displayed on User B's screen and played through the speakers.

[0149] In this way, the system of the present invention can provide high-quality video and audio in real time while significantly reducing the amount of data sent and received and easing the strain on network resources.

[0150] (Application example 1)

[0151] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0152] Conventional online conference systems and video streaming systems require large amounts of data for transmission and reception, resulting in problems such as strain on network resources and communication delays. Furthermore, real-time natural dialogue between customers and store staff in virtual stores requires the transmission and reception of high-quality video and audio. However, this data transmission depends on bandwidth and communication quality, which can potentially impair the customer experience. A system that can solve this issue and efficiently provide real-time video and audio is needed.

[0153] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0154] In this invention, the server includes a means for receiving and storing customer profile images and voice samples, a means for capturing video and audio of the virtual store clerk in real time, extracting and compressing feature data, and transmitting the data to the server, and a means for the receiving terminal to play back video and audio of the virtual store clerk in real time using artificial intelligence for generation, thereby enabling efficient use of network resources and providing high-quality real-time video and audio.

[0155] The "transmitting terminal" is a terminal that captures video and audio data and extracts feature data in real time.

[0156] "Feature data" is data extracted from video and audio data, such as facial landmarks and audio frequency spectrum characteristics.

[0157] The "means for compressing and forming transmission packets" refers to a means for efficiently compressing the extracted feature data and converting it into a packet format for transmission.

[0158] The "recipient's terminal" is a terminal that receives the feature data sent from the sender and generates video and audio in real time using artificial intelligence.

[0159] "Generative AI" is an AI technology that generates video and audio in real time based on feature data.

[0160] A "virtual store" is a store that exists in a virtual space and that users can access online to purchase goods and services.

[0161] A "profile picture" is a static image that a user provides when connecting to an online service.

[0162] "Audio sample" is audio data provided by a user when connecting to an online service.

[0163] A "virtual store clerk" is a virtual character that serves customers in a virtual store.

[0164] "Means for playing in real time" refers to means for instantly generating video and audio using artificial intelligence based on received feature data, and displaying and playing them to the user.

[0165] The system for implementing this invention includes a real-time customer service system in a virtual store. The system is composed of the following main components:

[0166] Server-side processing

[0167] When a customer visits the virtual store, the server receives and stores the customer's profile image and voice sample (e.g., self-introduction). The profile data is used for customer identification and data matching. Next, the server receives real-time video and audio data sent from the store clerk's terminal and extracts feature data from that data. Specifically, this includes facial landmarks and frequency spectrum characteristics of the voice. The extracted feature data is efficiently compressed and sent to the recipient's terminal.

[0168] Terminal side processing

[0169] The recipient's device (customer's device) receives the feature data sent from the server. Based on the received feature data, a generative AI model is used to generate video and audio of a virtual store clerk in real time. The generated video and audio are then played back to the customer in real time, allowing the customer to enjoy a natural communication experience, as if they were in a real store.

[0170] Hardware and software used

[0171] Server: A server with powerful CPU and GPU. Uses generative artificial intelligence models (e.g., OpenAI's GPT series).

[0172] Client device: Smartphone or head-mounted display (HMD). The OpenCV library is used to receive feature data and play back the generated video and audio.

[0173] Specific examples

[0174] Example prompt:

[0175] "When customers visit a virtual store, they communicate with a virtual associate in real time based on their visit."

[0176] This invention solves the data transmission and reception issues of conventional online conference systems and video streaming systems. It also enables real-time transmission of high-quality video and audio in virtual stores, improving the customer experience and enabling efficient use of network resources.

[0177] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0178] Step 1:

[0179] When a user visits the virtual store, the server receives and stores the user's profile picture and voice sample (e.g., self-introduction), which are used to identify the user and provide customized services.

[0180] Input: User's profile picture and voice sample.

[0181] Output: User profile data stored on the server.

[0182] How it works: The server stores user-submitted profile images and voice samples in a database.

[0183] Step 2:

[0184] The virtual store assistant uses a dedicated device to capture his or her own video and audio in real time, and the transmitting device extracts feature data such as facial landmarks and frequency spectrum characteristics from this data, then efficiently compresses and converts it into transmission packets.

[0185] Input: Real-time video and audio data.

[0186] Output: Outgoing packets containing compressed feature data.

[0187] How it works: The sending device analyzes the video and audio data to extract facial landmarks and audio frequency spectrum characteristics, then compresses and converts them into packets for transmission.

[0188] Step 3:

[0189] The server receives the transmitted packets from the sending terminal, analyzes them, and routes them appropriately to each receiving terminal.

[0190] Input: Transmission packets containing compressed feature data.

[0191] Output: Parsed feature data.

[0192] How it works: The server unpacks the received transmission packet and analyzes the feature data against the profile data stored in its internal database.

[0193] Step 4:

[0194] The receiver's device receives the feature data sent from the server and uses the generative AI model to generate real-time video and audio, which is then instantly played back to the user (customer).

[0195] Input: Feature data sent from the server.

[0196] Output: Real-time generated video and audio.

[0197] How it works: The recipient's device inputs the received feature data into a generative AI model, which then generates and plays back video and audio of a virtual store clerk based on this data.

[0198] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0199] The system of the present invention utilizes generative AI and an emotion engine to generate video and audio in real time and recognize emotions based on feature data in order to improve the efficiency of data transmission and reception in online meetings and video streaming.

[0200] Program processing overview

[0201] User connection and pre-transmission of data

[0202] The server receives profile images and voice samples of users when they connect to an online conference. During this process, users submit their own images and voice data introducing themselves, which are then stored on the server.

[0203] Extracting and sending feature data

[0204] The transmitting device captures the user's camera video and microphone audio in real time and extracts feature data from them, specifically facial landmarks (such as the positions of the eyes, nose, and mouth), gestures, and frequency spectrum characteristics from audio data.

[0205] The device recognizes the user's emotions based on the captured video and audio data using an emotion engine that analyzes emotions from facial expressions and voice tone.

[0206] The terminal compresses the extracted feature data and the recognized emotion data and transmits them to the server as efficient transmission packets.

[0207] Video Creation and Distribution

[0208] The server analyzes the feature and emotion data received from the sending device and transmits it to the receiving device in an appropriate format, allowing the generating AI to reproduce the video and audio in real time.

[0209] The receiving device uses the feature data and emotion data sent from the server to run the generative AI and generate video and audio in real time. The generated video and audio also reflect the recognized emotions, realizing natural communication.

[0210] Specific examples

[0211] 1. Connection and Pre-Data Transmission

[0212] User A connects to the online conference, and the server receives the profile picture and self-introduction voice of User A. User A's terminal transmits this data to the server.

[0213] 2. Extracting and sending feature data

[0214] The transmitting terminal (User A's device) captures the video and audio of User A in real time during the conference. From this data, feature data such as facial position information and audio frequency characteristics are extracted.

[0215] Furthermore, the emotion engine installed in User A's device recognizes User A's emotions from facial expressions and voice tone, analyzing emotions such as smiling or anger, for example.

[0216] The extracted feature data and emotion data are compressed and transmitted to the server as transmission packets.

[0217] 3. Video Creation and Distribution

[0218] The server analyzes the received feature data and emotion data and sends them in an appropriate format to the receiving device (User B's device). The receiving device inputs the transmitted data into the generation AI, which generates real-time video and audio of User A.

[0219] The generated video and audio also reflect User A's emotions, allowing User B to have a natural online meeting experience based on User A's facial expressions and tone.

[0220] In this way, the system of the present invention can provide natural communication through emotion recognition while significantly reducing the amount of data sent and received and easing the strain on network resources, resulting in cost reduction and a reduction in the environmental impact.

[0221] The processing flow will be explained below.

[0222] Step 1:

[0223] A user connects to an online meeting.

[0224] The user launches an online conference application and connects to the server by entering a URL or ID to connect to a specific conference room.

[0225] Step 2:

[0226] The server receives the advance data.

[0227] Upon initial connection, the server receives a profile image and voice sample from the user, including a photo of the user's face and a voice clip introducing themselves.

[0228] Step 3:

[0229] The user submits the advance data.

[0230] The user device sends a facial photo taken with a camera and a voice sample recorded with a microphone to the server, and this pre-data is used for user identification and initial setup.

[0231] Step 4:

[0232] The device captures video and audio.

[0233] The sending device captures the user's camera video and microphone audio in real time, recording the video and audio data frame by frame.

[0234] Step 5:

[0235] The device extracts the feature data.

[0236] The device detects facial landmarks (such as the eyes, nose, and mouth) and gestures from the captured video. Frequency spectrum characteristics are also extracted from the audio data. This allows the device to accurately understand the user's movements and speech while reducing the amount of data.

[0237] Step 6:

[0238] The device runs the emotion engine.

[0239] The device's built-in emotion engine recognizes the user's emotions from captured video and audio data, using a dedicated algorithm to identify emotions such as joy, anger, sadness, and happiness from facial expressions and voice tone.

[0240] Step 7:

[0241] The device compresses the feature data and emotion data.

[0242] The device efficiently compresses the extracted feature data and emotion data and sends them as packets, which are small in size and minimize network load.

[0243] Step 8:

[0244] The terminal transmits a transmission packet to the server.

[0245] The device sends packets containing compressed feature data and emotion data to the server in real time, a process that is fast and causes almost no delay.

[0246] Step 9:

[0247] The server analyzes the feature data and emotion data.

[0248] The server analyzes the received transmission packets and decodes the feature data and emotion data, converting them into a format that can be used by the receiving terminal.

[0249] Step 10:

[0250] The server transmits the data to the receiving terminal.

[0251] The server then transmits the analyzed feature data and emotion data to the receiving device, which receives this data and uses it as input for the generation AI.

[0252] Step 11:

[0253] The receiving device launches the generation AI.

[0254] The receiving device inputs the transmitted feature data and emotion data into the generation AI, initiating the real-time video and audio generation process.

[0255] Step 12:

[0256] Generative AI generates real-time video and audio.

[0257] The generative AI generates real-time video and audio of the user from feature data, and also reflects the user's facial expressions and vocal intonation based on emotional data.

[0258] Step 13:

[0259] The receiving terminal displays the generated video and audio to the user.

[0260] The receiving terminal displays the generated video and audio to the user in real time, allowing the user to view natural video and audio that reflects the emotions of the sending user.

[0261] This system significantly reduces the amount of data sent and received, alleviating the burden on network resources while providing natural communication through emotion recognition, resulting in cost savings and a lighter environmental impact.

[0262] Example 2

[0263] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0264] In online meetings and video streaming, it is important to generate high-quality video and audio in real time to improve the efficiency of data transmission and reception. It is also necessary to recognize user emotions and realize natural communication.

[0265] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a sending terminal including means for extracting feature data from real-time video and audio data, means for recognizing a user's emotion based on the feature data, means for compressing the feature data and emotion data into transmission packets, means for transmitting the transmission packets to a receiving terminal, and means for the receiving terminal to generate real-time video and audio using a generation AI based on the feature data and emotion data. This enables efficient data transmission and reception and natural communication.

[0266] A "transmitting terminal" is a device that captures video and audio data in real time, extracts feature data from the data, and transmits the data to a server.

[0267] "Feature data" refers to information such as facial landmarks and frequency spectrum characteristics extracted from the user's video and audio data.

[0268] "Emotion data" is information about emotions analyzed from the user's video and audio by an emotion recognition engine.

[0269] A "transmission packet" is a data packet that contains compressed feature data and emotion data.

[0270] The "recipient's terminal" is a device that receives the transmission packets sent from the server and generates video and audio in real time using generation AI.

[0271] "Generative AI" is an artificial intelligence model that generates video and audio in real time based on feature data and emotional data.

[0272] An "emotion recognition engine" is software that analyzes emotions from a user's video and audio data and outputs them as emotional data.

[0273] A "machine learning model" is a model that uses algorithms to learn patterns and regularities from data and make inferences and predictions.

[0274] The system of the present invention improves the efficiency of data transmission and reception in online meetings and video streaming, and utilizes generative AI and an emotion engine to generate video and audio in real time and recognize emotions based on feature data.

[0275] User connection and pre-transmission of data

[0276] When a user connects to an online conference, the user sends their profile picture and self-introduction audio data from their device to the server, which then receives and stores this data for each user.

[0277] Extracting and sending feature data

[0278] The sending device captures the user's camera video and microphone audio in real time during an online conference. From this captured data, the device extracts frequency spectrum characteristics from facial landmarks (such as the positions of the eyes, nose, and mouth), gestures, and voice data. Furthermore, it uses an emotion engine to recognize the user's emotions and generate emotion data.

[0279] The extracted feature data and emotion data are compressed and sent to the server as transmission packets. This compression and transmission process is important for enabling efficient communication.

[0280] Video Creation and Distribution

[0281] The server analyzes the transmitted feature data and emotion data and transmits it to the recipient's device in the appropriate format. This data is formatted so that the generative AI can reproduce the video and audio in real time. The recipient's device inputs the data transmitted from the server into the generative AI model, generating the video and audio in real time. The generated video and audio reflect the user's emotions, enabling natural communication.

[0282] Specific examples

[0283] 1. Connection and Pre-Data Transmission

[0284] When user A connects to an online conference, the terminal sends a profile image (profile.jpg) and self-introduction audio data (intro.mp3) to the server.

[0285] The server receives this data, associates it with User A's ID, and stores it in a database.

[0286] 2. Extracting and sending feature data

[0287] The transmitting device captures the video and audio of User A in real time. From this captured data, facial landmarks (eyes (x1, y1), nose (x2, y2), mouth (x3, y3)) and audio frequency characteristics (frequency band Hz, amplitude dB) are extracted.

[0288] The emotion engine recognizes emotions from user A's facial expressions and tone of voice, and generates emotion data for "joy," for example.

[0289] The extracted feature data and emotion data are compressed and transmitted to the server as transmission packets.

[0290] 3. Video Creation and Distribution

[0291] The server analyzes the received feature data and emotion data and transmits them to the receiving terminal (user B's terminal) in an appropriate format.

[0292] The receiving device inputs the transmitted feature data and emotion data into the generative AI model, generating video and audio of User A in real time. Because the generated video and audio also reflect User A's emotions, User B can enjoy a natural online meeting experience based on User A's facial expressions and tone.

[0293] Prompt Sentence Examples

[0294] "Analyze the camera video and microphone audio of User A, and generate video and audio in real time based on the following feature data. The feature data is as follows:

[0295] Facial landmarks: eyes(x,y), nose(x,y), mouth(x,y)

[0296] Audio frequency spectrum: frequency band (Hz), amplitude (dB)

[0297] Emotion data: smile, anger

[0298] The generated video and audio should reflect the perceived emotions of User A.

[0299] In this way, the system of the present invention can provide efficient data transmission and reception and natural communication, thereby reducing costs and environmental impact.

[0300] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0301] Step 1:

[0302] When a user connects to an online conference, the user sends a profile image and self-introduction voice data from the terminal to the server. The server receives the user's profile image and self-introduction voice data and stores them in a database for each user.

[0303] Input: User's profile image (e.g., profile.jpg) and self-introduction audio data (e.g., intro.mp3)

[0304] Output: User profile image and self-introduction audio data stored in the server database

[0305] Specific operation: The user selects a profile picture and voice data on the device and clicks the "Send" button. The server associates the received data with the user ID and stores it in the database.

[0306] Step 2:

[0307] The sending device captures the user's camera video and microphone audio in real time during the online conference, so that the user's current video and audio are input to the device.

[0308] Input: Real-time captured video and audio data

[0309] Output: Video frames and audio samples captured in real time

[0310] Specific operation: The camera captures the user's image and the microphone records the user's voice. These data are stored in the device's memory.

[0311] Step 3:

[0312] The transmitting terminal extracts facial landmarks (such as the positions of the eyes, nose, and mouth) and gestures from the captured video data, and frequency spectrum characteristics from the audio data.

[0313] Input: Video frames and audio samples captured in real time

[0314] Output: Extracted facial landmarks (e.g., eye position (x1, y1), nose position (x2, y2), mouth position (x3, y3)), frequency spectrum characteristics (e.g., frequency band Hz, amplitude dB)

[0315] How it works: The image processing algorithm analyzes the video data and detects facial features, while the audio analysis algorithm analyzes the frequency characteristics of the audio data.

[0316] Step 4:

[0317] The transmitting terminal uses an emotion engine to recognize emotions from the user's facial expressions and tone of voice, and generates emotion data.

[0318] Input: Video frames and audio samples captured in real time

[0319] Output: Recognized emotion data (e.g., "joy", "anger")

[0320] Specific operation: The emotion recognition engine analyzes facial expressions and voice to identify the user's emotional state, resulting in emotional data.

[0321] Step 5:

[0322] The transmitting terminal compresses the extracted feature data and emotion data and transmits them as transmission packets to the server.

[0323] Input: extracted feature data and emotion data

[0324] Output: Compressed outgoing packets

[0325] Specific operation: The data compression algorithm compresses the feature data and emotion data to generate a transmission packet, which the device then sends to the server.

[0326] Step 6:

[0327] The server analyzes the transmitted feature data and emotion data and transmits them to the receiving terminal in an appropriate format.

[0328] Input: Outgoing packets

[0329] Output: Data reformatted to the appropriate format

[0330] Specific operation: The server decompresses and analyzes the transmitted packets, reformats the analysis results into a format usable by the receiving terminal, and transmits them.

[0331] Step 7:

[0332] The receiving terminal inputs the feature data and emotion data sent from the server into the generative AI model and generates video and audio in real time.

[0333] Input: Feature data and emotion data sent from the server

[0334] Output: Generated real-time video and audio

[0335] How it works: A prompt is input into the generative AI model, which generates video and audio in real time, with the recognized emotion reflected in the generated video and audio.

[0336] The above processing steps enable efficient data transmission and reception and natural communication in online conferences and video streaming.

[0337] (Application example 2)

[0338] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0339] Online conferences and video streaming services that utilize real-time video and audio data require the transmission and reception of large amounts of data, resulting in increased communication costs and network latency. Furthermore, it is difficult to accurately convey users' emotions, which can impede natural communication. This creates a demand for more efficient and natural communication methods.

[0340] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0341] In this invention, the server includes a sending terminal that extracts feature data from real-time video and audio data, a means for compressing the feature data and emotion data into transmission packets, a means for transmitting the transmission packets to a receiving terminal, and a means for the receiving terminal to generate real-time video and audio using a generation AI based on the feature data and emotion data. This makes it possible to generate natural video and audio that reflects the user's emotions while significantly reducing the amount of data sent and received.

[0342] The "transmitting terminal" is a device where a user generates video and audio data in real time and extracts this as feature data and emotion data.

[0343] "Feature data" is analytical information such as facial landmarks and audio frequency characteristics extracted from real-time video and audio data.

[0344] "Emotion data" is information obtained as a result of analyzing the user's emotional state based on the extracted feature data.

[0345] A "transmission packet" is a data packet containing compressed feature data and emotion data, which is transmitted over a network.

[0346] The "recipient terminal" is a device that receives the transmitted packets and runs the generation AI to generate video and audio in real time.

[0347] "Generative AI" is an artificial intelligence technology that generates video and audio in real time based on received feature data and emotional data.

[0348] A "specialized algorithm" is a program that uses specific processing techniques to extract feature data and emotion data from video and audio data.

[0349] A "machine learning model" is a trained model that, as part of generative AI, generates video and audio in real time based on incoming data.

[0350] MODE FOR CARRYING OUT THE INVENTION

[0351] An object of the present invention is to provide a system that improves the efficiency of data transmission and reception in online conferences and live streaming services, and realizes natural communication that reflects the user's emotions. The following describes in detail an embodiment of the present invention.

[0352] Program processing overview

[0353] The system of the present invention consists of three main components: a sending terminal, a server, and a receiving terminal.

[0354] Sending device

[0355] The transmitting device captures real-time video and audio and extracts feature and emotion data from it. Specifically, it acquires video and audio using hardware such as a camera and microphone, and analyzes facial landmarks and audio frequency characteristics using a dedicated algorithm. Based on the results of this analysis, the emotion engine recognizes the user's emotions. The extracted feature and emotion data are compressed and sent to the server as a transmission packet.

[0356] server

[0357] The server manages and analyzes the feature data and emotion data received from the sending device and sends it to the receiving device. To ensure efficient data transmission and reception, the server converts the feature data and emotion data into an appropriate format and prepares it in a form that is easy for the generation AI to process before sending it.

[0358] Recipient's device

[0359] The receiver's device runs a generation AI based on the data received from the server to generate real-time video and audio. The generation AI uses a machine learning model to reproduce the feature data and emotion data compressed for transmission in real time, generating natural-looking video and audio that also reflects the user's emotions.

[0360] Hardware and Software Used

[0361] Sending device: camera, microphone, specialized algorithms (e.g., facial landmark detection algorithms)

[0362] Emotion engine: Software for recognizing emotions from audio and video (e.g., EmotionEngine)

[0363] Server: Server software that manages and analyzes data transmission and reception

[0364] Recipient device: Generative AI model (e.g., a generative AI using a specific machine learning model)

[0365] Data processing and calculation

[0366] Processing on the sending device: Facial landmarks and audio frequency characteristics are extracted from video and audio captured in real time, and emotions are recognized using an emotion engine.

[0367] Data compression and transmission: The extracted feature data and emotion data are compressed and sent to the server as a transmission packet.

[0368] Processing on the server: The received data is analyzed, converted into a format that is easy for the generating AI to process, and sent to the recipient's device.

[0369] Processing on the recipient's device: Using generative AI, video and audio are generated in real time based on the transmitted feature data and emotion data.

[0370] Specific examples

[0371] When a live streamer broadcasts live on their smartphone, they register a profile picture and audio sample in advance. During the live broadcast, the sending device captures video and audio in real time, extracts feature data and emotional data, and sends it to a server. The receiving device runs a generative AI based on the received data to generate video and audio in real time. This method makes it possible to provide viewers with high-quality video and audio with low latency.

[0372] Example prompt sentence:

[0373] "Generate video and audio in real time based on the user's profile picture and voice sample. Recognize the user's emotions from facial expressions and tone of voice."

[0374] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0375] Step 1:

[0376] The sending device captures video and audio in real time from the user's camera and microphone. The input is the camera video and microphone audio, and the output is the captured video frames and audio samples of these data. Specifically, it performs a process to acquire the data stream from the camera and microphone.

[0377] Step 2:

[0378] The transmitting device detects facial landmarks (such as the eyes, nose, and mouth) from the captured video and extracts frequency spectrum characteristics from the audio. The input is video frames and audio samples, and the output is facial landmark data and audio characteristic data. Specifically, it runs a facial landmark detection algorithm and an audio spectrum analysis algorithm.

[0379] Step 3:

[0380] The transmitting device uses an emotion engine to recognize the user's emotion based on the detected facial landmarks and voice characteristic data. The input is facial landmark data and voice characteristic data, and the output is recognized emotion data. Specifically, the emotion engine is executed to analyze the user's emotion from their facial expression and voice tone.

[0381] Step 4:

[0382] The transmitting terminal compresses the extracted feature data and emotion data into transmission packets. The input is facial landmark data, voice characteristic data, and emotion data, and the output is compressed transmission packets. Specifically, the data is efficiently packetized using a data compression algorithm.

[0383] Step 5:

[0384] The sending terminal sends compressed transmission packets to the server. The input is the transmission packet, and the output is the data sent to the server. Specifically, the data is sent to the server using a network protocol.

[0385] Step 6:

[0386] The server decompresses the received transmission packets and analyzes the feature data and emotion data. The input is the received transmission packets, and the output is the decompressed feature data and emotion data. Specifically, the server executes a process to decompress the packets using a data decompression algorithm.

[0387] Step 7:

[0388] The server transmits the feature data and emotion data to the recipient's terminal. The input is the decompressed feature data and emotion data, and the output is the data transmitted to the recipient's terminal. Specifically, the data is transmitted using a network protocol.

[0389] Step 8:

[0390] The receiver's device inputs the received feature data and emotion data into the generative AI to generate real-time video and audio. The input is the received feature data and emotion data, and the output is the generated video and audio. Specifically, the generative AI model is used to process the video and audio based on the feature data and emotion data.

[0391] Step 9:

[0392] The receiver's terminal presents the generated real-time video and audio to the user. The input is the generated video and audio, and the output is the video and audio presented to the user. Specifically, the system performs a process to play back the generated content using a display and speakers.

[0393] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0394] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0395] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0396] [Second embodiment]

[0397] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0398] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0399] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0400] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0401] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0402] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0403] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0404] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0405] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0406] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0407] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0408] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0409] The system of the present invention utilizes generative AI to generate video and audio in real time based on feature data in order to improve the efficiency of data transmission and reception in online conferences and video streaming.

[0410] Program processing overview

[0411] User connection and pre-transmission of data

[0412] The server receives profile images and voice samples of users when they connect to an online conference. During this process, users submit their own images and voice data introducing themselves, which are then stored on the server.

[0413] Extracting and sending feature data

[0414] The transmitting device captures the user's camera video and microphone audio in real time and extracts feature data from them, including facial landmarks (such as the positions of the eyes, nose, and mouth) and frequency spectrum characteristics of the voice.

[0415] The terminal compresses the extracted feature data and transmits it to the server as an efficient transmission packet.

[0416] Video Creation and Distribution

[0417] The server analyzes the feature data received from the sending device and transmits it to the receiving device in an appropriate format, allowing the generating AI to recreate the video and audio in real time.

[0418] The receiving device runs the generative AI using the feature data sent from the server to generate video and audio in real time, which is then displayed to the receiving user, providing an experience similar to that of a regular online meeting.

[0419] Specific examples

[0420] 1. Connection and Pre-Data Transmission

[0421] User A connects to the online conference, and the server receives the profile picture and self-introduction voice of User A. User A's terminal transmits this data to the server.

[0422] 2. Extracting and sending feature data

[0423] The transmitting terminal (User A's device) captures User A's video and audio in real time during the conference and extracts feature data such as facial position information and audio frequency characteristics. This feature data is efficiently compressed and sent to the server as a transmission packet.

[0424] 3. Video Creation and Distribution

[0425] The server analyzes the received feature data and sends it in an appropriate format to the receiving device (User B's device). The receiving device inputs the transmitted feature data into the generation AI, which generates real-time video and audio of User A. User B then watches the generated video and audio, enjoying a natural online meeting experience.

[0426] In this way, the system of the present invention can provide real-time video and audio while significantly reducing the amount of data transmitted and received and easing the strain on network resources, resulting in cost reduction and a reduction in the environmental impact.

[0427] The processing flow will be explained below.

[0428] Step 1:

[0429] A user connects to an online meeting.

[0430] The user launches the online meeting application and connects to the server.

[0431] Step 2:

[0432] The server receives the advance data.

[0433] The server receives a profile image and voice sample from the user, which includes a photo of the user's face and an audio introduction.

[0434] Step 3:

[0435] The user submits the advance data.

[0436] Users send their profile picture and voice sample from their device to the server, and this data is sent when the meeting is initially connected.

[0437] Step 4:

[0438] The device captures video and audio.

[0439] The sending device captures the user's camera video and microphone audio in real time, allowing conversations and actions to be recorded.

[0440] Step 5:

[0441] The device extracts the feature data.

[0442] The device extracts facial landmarks (such as the positions of the eyes, nose, and mouth) and gestures from the captured video, and also extracts frequency spectrum characteristics from the audio data.

[0443] Step 6:

[0444] The terminal compresses the feature data.

[0445] The terminal efficiently compresses the extracted feature data to generate transmission packets, which significantly reduces the amount of data.

[0446] Step 7:

[0447] The terminal transmits the feature data to the server.

[0448] The terminal transmits a transmission packet to the server in real time, and the transmission packet includes compressed feature data.

[0449] Step 8:

[0450] The server analyzes the feature data.

[0451] The server parses the received feature data and decodes it into the appropriate format, making the data usable by the receiving device.

[0452] Step 9:

[0453] The server transmits the data to the receiving terminal.

[0454] The server transmits the analyzed feature data to the receiving terminal, which then generates the data in real time.

[0455] Step 10:

[0456] The receiving device launches the generation AI.

[0457] The receiving device activates the generation AI based on the transmitted feature data and inputs the data, which starts the generation process.

[0458] Step 11:

[0459] Generative AI generates real-time video and audio.

[0460] The generative AI uses the feature data to generate real-time video and audio that replicates the movements and mouth movements of the sending user.

[0461] Step 12:

[0462] The receiving terminal displays the generated video and audio to the user.

[0463] The receiving terminal displays the generated video and audio to the user in real time, and the user can see and hear the video and audio of other participants just like in a regular online conference.

[0464] This system significantly reduces the amount of data sent and received, easing the load on network resources and resulting in cost savings and a reduced environmental impact.

[0465] Example 1

[0466] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0467] In online conferences and video streaming, sending and receiving real-time video and audio data requires a large amount of network resources, resulting in communication delays and quality degradation. Furthermore, existing technologies lack efficient methods for generating high-quality video and audio in real time. This results in poor user experience and increases costs and environmental impact.

[0468] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0469] In this invention, the server includes a means for a user to connect to an online conference and transmit a profile image and a voice sample, a transmitting terminal including means for extracting feature data including facial landmarks and voice frequency spectrum characteristics from real-time video and audio data, a means for compressing the feature data into a transmission packet, a means for transmitting the transmission packet to the server and for the server to convert the feature data into an appropriate format, a means for the server to transmit the converted data to a receiving terminal, and a means for the receiving terminal to generate real-time video and audio using a generative AI model based on the feature data. This significantly reduces the amount of data transmitted and received, enabling real-time generation of high-quality video and audio while minimizing communication delays.

[0470] "Users" are typical participants who use online meetings and video streaming services.

[0471] An "online conference" refers to a conference or meeting that takes place over the Internet in real time, sharing audio and video.

[0472] A "profile image" is a still image that a user uses to indicate their personal information or identity.

[0473] A "voice sample" is a short piece of voice data that a user provides to the server, used for self-introduction and identification.

[0474] The "transmitting terminal" is a device used by a user (for example, a smartphone or a personal computer) that captures video and audio using a camera and microphone.

[0475] "Feature data" are specific digital indices or characteristics (e.g., facial landmarks or frequency spectrum characteristics of voice) extracted from captured video and audio data.

[0476] "Landmarks" are characteristic points that indicate the position information of each part of the face (eyes, nose, mouth, etc.).

[0477] The "frequency spectrum characteristics" are data that indicate the strength of each frequency component in an audio signal.

[0478] A "server" is a central system that manages and processes data sent and received from sending and receiving terminals.

[0479] A "transmission packet" is a small unit of compressed data containing feature data that is transmitted over a network.

[0480] The "receiving terminal" is a device that receives the feature data sent from the server and reproduces real-time video and audio using a generation AI.

[0481] A "generative AI model" is an artificial intelligence technology that uses machine learning algorithms to generate video and audio from feature data.

[0482] "Compression" is a technique for reducing data size and is used to improve transmission efficiency.

[0483] "Real-time" refers to instantaneous processing and communication with little or no delay.

[0484] MODE FOR CARRYING OUT THE INVENTION

[0485] The system of the present invention utilizes a generative AI model to generate video and audio in real time based on feature data in order to improve the efficiency of data transmission and reception in online meetings and video streaming. A specific embodiment of this system is described below.

[0486] Hardware and software used

[0487] Hardware: A computer device with a compatible camera, microphone, and internet connection

[0488] Software: Generative AI models (e.g., GPT-3, DALL-E), libraries for video and audio capture (e.g., OpenCV, PyAudio)

[0489] A natural language description of the program

[0490] User connection and pre-transmission of data

[0491] A user opens an online conference app on their computer and connects to the server. The user then sends a profile picture and a self-introduction voice sample. During this process, the user follows the application's instructions to select a profile picture they have taken in advance, record a self-introduction voice using a microphone, and send it to the server. The server then stores the received data.

[0492] Real-time extraction and transmission of feature data

[0493] The sending device (e.g., User A's device) uses a camera and microphone to capture User A's video and audio in real time. Specifically, it uses a computer vision library such as OpenCV to detect facial landmarks and an audio analysis library such as PyAudio to analyze the frequency spectrum characteristics of the audio. These feature data are efficiently compressed and sent to the server as transmission packets.

[0494] Feature data analysis and resubmission

[0495] The server analyzes the feature data received from the sending device and converts the data into the appropriate format. This involves using generative AI models to calculate additional data needed to generate the video and audio. For example, it generates data to complete facial details and audio. The server then transmits the converted data to the receiving device.

[0496] Real-time video and audio generation

[0497] The receiving device (e.g., User B's device) receives the feature data and complementary data sent from the server, and the internal generative AI model is activated. For example, DALL-E generates a video reproducing User A's face based on the facial landmark data and complementary data, and simultaneously generates real-time audio based on the audio sample. The generated video and audio are output to the screen and speakers, just like a regular video conferencing app.

[0498] Specific examples

[0499] 1. Connection and Pre-Data Transmission

[0500] User A connects to the online conference, and the server receives the profile picture and self-introduction voice of User A. User A's terminal transmits this data to the server.

[0501] 2. Real-time extraction and transmission of feature data

[0502] The transmitting terminal (User A's device) captures User A's video and audio in real time during the conference and extracts feature data such as facial position information and audio frequency characteristics. This feature data is efficiently compressed and sent to the server as a transmission packet.

[0503] 3. Analysis and retransmission of feature data

[0504] The server analyzes the received feature data and sends it in an appropriate format to the receiving device of User B. User B's device inputs the transmitted feature data into the generation AI, which generates real-time video and audio of User A.

[0505] Prompt Sentence Examples

[0506] Data sent from User A's device to the server: "Send profile picture and voice sample"

[0507] The process in which the server sends feature data to User B's device: "Generate video and audio based on facial landmarks and audio spectrum characteristics."

[0508] This system significantly reduces the amount of data transmitted and received, alleviating pressure on network resources while providing high-quality real-time video and audio, resulting in cost savings and a lighter environmental impact.

[0509] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0510] Step 1:

[0511] User connection and pre-transmission of data

[0512] A user opens an online conference app on their computer and connects to the server. The user provides a profile image and a self-introduction voice sample as input, which is then sent to the server. The input image and voice data are stored on the server. For example, the user may upload a face photo taken with a camera or a voice recording with a microphone. This process is performed to collect initial data that will serve as the basis for subsequent feature data generation.

[0513] Step 2:

[0514] Real-time extraction and transmission of feature data

[0515] The sending terminal (user's device) uses a camera and microphone to capture the user's video and audio in real time. This becomes the input data. Specifically, a computer vision library such as OpenCV is used to detect facial landmarks (the positions of the eyes, nose, mouth, etc.), and an audio analysis library such as PyAudio is used to analyze the frequency spectrum characteristics of the audio. This feature data is compressed and sent to the server as a transmission packet. For example, this includes operations to obtain the position information of each part of the face from the camera video and calculate the strength of specific frequency components from the audio signal. The output is compressed feature data.

[0516] Step 3:

[0517] Feature data analysis and resubmission

[0518] The server receives feature data from the sending device as input. It analyzes this data and uses a generative AI model (e.g., GPT-3 or DALL-E) to calculate the necessary completion data for video and audio generation. Specifically, it generates completion data for facial details and audio. This data is then compressed again and sent to the receiving device in an appropriate format. For example, this may involve the generative AI completing fine details of the user's face to generate high-quality video data. The output is a packet containing the completion data and the original feature data.

[0519] Step 4:

[0520] Real-time video and audio generation

[0521] The receiving device (User B's device) receives the feature data and complementary data from the server as input. An internal generative AI model is activated and generates video and audio in real time based on this data. Specifically, DALL-E generates video recreating User A's face based on facial landmark data, and generates real-time audio based on audio samples. The output is the generated video and audio, which are displayed on User B's screen and played through the speakers.

[0522] In this way, the system of the present invention can provide high-quality video and audio in real time while significantly reducing the amount of data sent and received and easing the strain on network resources.

[0523] (Application example 1)

[0524] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0525] Conventional online conference systems and video streaming systems require large amounts of data for transmission and reception, resulting in problems such as strain on network resources and communication delays. Furthermore, real-time natural dialogue between customers and store staff in virtual stores requires the transmission and reception of high-quality video and audio. However, this data transmission depends on bandwidth and communication quality, which can potentially impair the customer experience. A system that can solve this issue and efficiently provide real-time video and audio is needed.

[0526] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0527] In this invention, the server includes a means for receiving and storing customer profile images and voice samples, a means for capturing video and audio of the virtual store clerk in real time, extracting and compressing feature data, and transmitting the data to the server, and a means for the receiving terminal to play back video and audio of the virtual store clerk in real time using artificial intelligence for generation, thereby enabling efficient use of network resources and providing high-quality real-time video and audio.

[0528] The "transmitting terminal" is a terminal that captures video and audio data and extracts feature data in real time.

[0529] "Feature data" is data extracted from video and audio data, such as facial landmarks and audio frequency spectrum characteristics.

[0530] The "means for compressing and forming transmission packets" refers to a means for efficiently compressing the extracted feature data and converting it into a packet format for transmission.

[0531] The "recipient's terminal" is a terminal that receives the feature data sent from the sender and generates video and audio in real time using artificial intelligence.

[0532] "Generative AI" is an AI technology that generates video and audio in real time based on feature data.

[0533] A "virtual store" is a store that exists in a virtual space and that users can access online to purchase goods and services.

[0534] A "profile picture" is a static image that a user provides when connecting to an online service.

[0535] "Audio sample" is audio data provided by a user when connecting to an online service.

[0536] A "virtual store clerk" is a virtual character that serves customers in a virtual store.

[0537] "Means for playing in real time" refers to means for instantly generating video and audio using artificial intelligence based on received feature data, and displaying and playing them to the user.

[0538] The system for implementing this invention includes a real-time customer service system in a virtual store. The system is composed of the following main components:

[0539] Server-side processing

[0540] When a customer visits the virtual store, the server receives and stores the customer's profile image and voice sample (e.g., self-introduction). The profile data is used for customer identification and data matching. Next, the server receives real-time video and audio data sent from the store clerk's terminal and extracts feature data from that data. Specifically, this includes facial landmarks and frequency spectrum characteristics of the voice. The extracted feature data is efficiently compressed and sent to the recipient's terminal.

[0541] Terminal side processing

[0542] The recipient's device (customer's device) receives the feature data sent from the server. Based on the received feature data, a generative AI model is used to generate video and audio of a virtual store clerk in real time. The generated video and audio are then played back to the customer in real time, allowing the customer to enjoy a natural communication experience, as if they were in a real store.

[0543] Hardware and software used

[0544] Server: A server with powerful CPU and GPU. Uses generative artificial intelligence models (e.g., OpenAI's GPT series).

[0545] Client device: Smartphone or head-mounted display (HMD). The OpenCV library is used to receive feature data and play back the generated video and audio.

[0546] Specific examples

[0547] Example prompt:

[0548] "When customers visit a virtual store, they communicate with a virtual associate in real time based on their visit."

[0549] This invention solves the data transmission and reception issues of conventional online conference systems and video streaming systems. It also enables real-time transmission of high-quality video and audio in virtual stores, improving the customer experience and enabling efficient use of network resources.

[0550] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0551] Step 1:

[0552] When a user visits the virtual store, the server receives and stores the user's profile picture and voice sample (e.g., self-introduction), which are used to identify the user and provide customized services.

[0553] Input: User's profile picture and voice sample.

[0554] Output: User profile data stored on the server.

[0555] How it works: The server stores user-submitted profile images and voice samples in a database.

[0556] Step 2:

[0557] The virtual store assistant uses a dedicated device to capture his or her own video and audio in real time, and the transmitting device extracts feature data such as facial landmarks and frequency spectrum characteristics from this data, then efficiently compresses and converts it into transmission packets.

[0558] Input: Real-time video and audio data.

[0559] Output: Outgoing packets containing compressed feature data.

[0560] How it works: The sending device analyzes the video and audio data to extract facial landmarks and audio frequency spectrum characteristics, then compresses and converts them into packets for transmission.

[0561] Step 3:

[0562] The server receives the transmitted packets from the sending terminal, analyzes them, and routes them appropriately to each receiving terminal.

[0563] Input: Transmission packets containing compressed feature data.

[0564] Output: Parsed feature data.

[0565] How it works: The server unpacks the received transmission packet and analyzes the feature data against the profile data stored in its internal database.

[0566] Step 4:

[0567] The receiver's device receives the feature data sent from the server and uses the generative AI model to generate real-time video and audio, which is then instantly played back to the user (customer).

[0568] Input: Feature data sent from the server.

[0569] Output: Real-time generated video and audio.

[0570] How it works: The recipient's device inputs the received feature data into a generative AI model, which then generates and plays back video and audio of a virtual store clerk based on this data.

[0571] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0572] The system of the present invention utilizes generative AI and an emotion engine to generate video and audio in real time and recognize emotions based on feature data in order to improve the efficiency of data transmission and reception in online meetings and video streaming.

[0573] Program processing overview

[0574] User connection and pre-transmission of data

[0575] The server receives profile images and voice samples of users when they connect to an online conference. During this process, users submit their own images and voice data introducing themselves, which are then stored on the server.

[0576] Extracting and sending feature data

[0577] The transmitting device captures the user's camera video and microphone audio in real time and extracts feature data from them, specifically facial landmarks (such as the positions of the eyes, nose, and mouth), gestures, and frequency spectrum characteristics from audio data.

[0578] The device recognizes the user's emotions based on the captured video and audio data using an emotion engine that analyzes emotions from facial expressions and voice tone.

[0579] The terminal compresses the extracted feature data and the recognized emotion data and transmits them to the server as efficient transmission packets.

[0580] Video Creation and Distribution

[0581] The server analyzes the feature and emotion data received from the sending device and transmits it to the receiving device in an appropriate format, allowing the generating AI to reproduce the video and audio in real time.

[0582] The receiving device uses the feature data and emotion data sent from the server to run the generative AI and generate video and audio in real time. The generated video and audio also reflect the recognized emotions, realizing natural communication.

[0583] Specific examples

[0584] 1. Connection and Pre-Data Transmission

[0585] User A connects to the online conference, and the server receives the profile picture and self-introduction voice of User A. User A's terminal transmits this data to the server.

[0586] 2. Extracting and sending feature data

[0587] The transmitting terminal (User A's device) captures the video and audio of User A in real time during the conference. From this data, feature data such as facial position information and audio frequency characteristics are extracted.

[0588] Furthermore, the emotion engine installed in User A's device recognizes User A's emotions from facial expressions and voice tone, analyzing emotions such as smiling or anger, for example.

[0589] The extracted feature data and emotion data are compressed and transmitted to the server as transmission packets.

[0590] 3. Video Creation and Distribution

[0591] The server analyzes the received feature data and emotion data and sends them in an appropriate format to the receiving device (User B's device). The receiving device inputs the transmitted data into the generation AI, which generates real-time video and audio of User A.

[0592] The generated video and audio also reflect User A's emotions, allowing User B to have a natural online meeting experience based on User A's facial expressions and tone.

[0593] In this way, the system of the present invention can provide natural communication through emotion recognition while significantly reducing the amount of data sent and received and easing the strain on network resources, resulting in cost reduction and a reduction in the environmental impact.

[0594] The processing flow will be explained below.

[0595] Step 1:

[0596] A user connects to an online meeting.

[0597] The user launches an online conference application and connects to the server by entering a URL or ID to connect to a specific conference room.

[0598] Step 2:

[0599] The server receives the advance data.

[0600] Upon initial connection, the server receives a profile image and voice sample from the user, including a photo of the user's face and a voice clip introducing themselves.

[0601] Step 3:

[0602] The user submits the advance data.

[0603] The user device sends a facial photo taken with a camera and a voice sample recorded with a microphone to the server, and this pre-data is used for user identification and initial setup.

[0604] Step 4:

[0605] The device captures video and audio.

[0606] The sending device captures the user's camera video and microphone audio in real time, recording the video and audio data frame by frame.

[0607] Step 5:

[0608] The device extracts the feature data.

[0609] The device detects facial landmarks (such as the eyes, nose, and mouth) and gestures from the captured video. Frequency spectrum characteristics are also extracted from the audio data. This allows the device to accurately understand the user's movements and speech while reducing the amount of data.

[0610] Step 6:

[0611] The device runs the emotion engine.

[0612] The device's built-in emotion engine recognizes the user's emotions from captured video and audio data, using a dedicated algorithm to identify emotions such as joy, anger, sadness, and happiness from facial expressions and voice tone.

[0613] Step 7:

[0614] The device compresses the feature data and emotion data.

[0615] The device efficiently compresses the extracted feature data and emotion data and sends them as packets, which are small in size and minimize network load.

[0616] Step 8:

[0617] The terminal transmits a transmission packet to the server.

[0618] The device sends packets containing compressed feature data and emotion data to the server in real time, a process that is fast and causes almost no delay.

[0619] Step 9:

[0620] The server analyzes the feature data and emotion data.

[0621] The server analyzes the received transmission packets and decodes the feature data and emotion data, converting them into a format that can be used by the receiving terminal.

[0622] Step 10:

[0623] The server transmits the data to the receiving terminal.

[0624] The server then transmits the analyzed feature data and emotion data to the receiving device, which receives this data and uses it as input for the generation AI.

[0625] Step 11:

[0626] The receiving device launches the generation AI.

[0627] The receiving device inputs the transmitted feature data and emotion data into the generation AI, initiating the real-time video and audio generation process.

[0628] Step 12:

[0629] Generative AI generates real-time video and audio.

[0630] The generative AI generates real-time video and audio of the user from feature data, and also reflects the user's facial expressions and vocal intonation based on emotional data.

[0631] Step 13:

[0632] The receiving terminal displays the generated video and audio to the user.

[0633] The receiving terminal displays the generated video and audio to the user in real time, allowing the user to view natural video and audio that reflects the emotions of the sending user.

[0634] This system significantly reduces the amount of data sent and received, alleviating the burden on network resources while providing natural communication through emotion recognition, resulting in cost savings and a lighter environmental impact.

[0635] Example 2

[0636] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0637] In online meetings and video streaming, it is important to generate high-quality video and audio in real time to improve the efficiency of data transmission and reception. It is also necessary to recognize user emotions and realize natural communication.

[0638] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a sending terminal including means for extracting feature data from real-time video and audio data, means for recognizing a user's emotion based on the feature data, means for compressing the feature data and emotion data into transmission packets, means for transmitting the transmission packets to a receiving terminal, and means for the receiving terminal to generate real-time video and audio using a generation AI based on the feature data and emotion data. This enables efficient data transmission and reception and natural communication.

[0639] A "transmitting terminal" is a device that captures video and audio data in real time, extracts feature data from the data, and transmits the data to a server.

[0640] "Feature data" refers to information such as facial landmarks and frequency spectrum characteristics extracted from the user's video and audio data.

[0641] "Emotion data" is information about emotions analyzed from the user's video and audio by an emotion recognition engine.

[0642] A "transmission packet" is a data packet that contains compressed feature data and emotion data.

[0643] The "recipient's terminal" is a device that receives the transmission packets sent from the server and generates video and audio in real time using generation AI.

[0644] "Generative AI" is an artificial intelligence model that generates video and audio in real time based on feature data and emotional data.

[0645] An "emotion recognition engine" is software that analyzes emotions from a user's video and audio data and outputs them as emotional data.

[0646] A "machine learning model" is a model that uses algorithms to learn patterns and regularities from data and make inferences and predictions.

[0647] The system of the present invention improves the efficiency of data transmission and reception in online meetings and video streaming, and utilizes generative AI and an emotion engine to generate video and audio in real time and recognize emotions based on feature data.

[0648] User connection and pre-transmission of data

[0649] When a user connects to an online conference, the user sends their profile picture and self-introduction audio data from their device to the server, which then receives and stores this data for each user.

[0650] Extracting and sending feature data

[0651] The sending device captures the user's camera video and microphone audio in real time during an online conference. From this captured data, the device extracts frequency spectrum characteristics from facial landmarks (such as the positions of the eyes, nose, and mouth), gestures, and voice data. Furthermore, it uses an emotion engine to recognize the user's emotions and generate emotion data.

[0652] The extracted feature data and emotion data are compressed and sent to the server as transmission packets. This compression and transmission process is important for enabling efficient communication.

[0653] Video Creation and Distribution

[0654] The server analyzes the transmitted feature data and emotion data and transmits it to the recipient's device in the appropriate format. This data is formatted so that the generative AI can reproduce the video and audio in real time. The recipient's device inputs the data transmitted from the server into the generative AI model, generating the video and audio in real time. The generated video and audio reflect the user's emotions, enabling natural communication.

[0655] Specific examples

[0656] 1. Connection and Pre-Data Transmission

[0657] When user A connects to an online conference, the terminal sends a profile image (profile.jpg) and self-introduction audio data (intro.mp3) to the server.

[0658] The server receives this data, associates it with User A's ID, and stores it in a database.

[0659] 2. Extracting and sending feature data

[0660] The transmitting device captures the video and audio of User A in real time. From this captured data, facial landmarks (eyes (x1, y1), nose (x2, y2), mouth (x3, y3)) and audio frequency characteristics (frequency band Hz, amplitude dB) are extracted.

[0661] The emotion engine recognizes emotions from user A's facial expressions and tone of voice, and generates emotion data for "joy," for example.

[0662] The extracted feature data and emotion data are compressed and transmitted to the server as transmission packets.

[0663] 3. Video Creation and Distribution

[0664] The server analyzes the received feature data and emotion data and transmits them to the receiving terminal (user B's terminal) in an appropriate format.

[0665] The receiving device inputs the transmitted feature data and emotion data into the generative AI model, generating video and audio of User A in real time. Because the generated video and audio also reflect User A's emotions, User B can enjoy a natural online meeting experience based on User A's facial expressions and tone.

[0666] Prompt Sentence Examples

[0667] "Analyze the camera video and microphone audio of User A, and generate video and audio in real time based on the following feature data. The feature data is as follows:

[0668] Facial landmarks: eyes(x,y), nose(x,y), mouth(x,y)

[0669] Audio frequency spectrum: frequency band (Hz), amplitude (dB)

[0670] Emotion data: smile, anger

[0671] The generated video and audio should reflect the perceived emotions of User A.

[0672] In this way, the system of the present invention can provide efficient data transmission and reception and natural communication, thereby reducing costs and environmental impact.

[0673] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0674] Step 1:

[0675] When a user connects to an online conference, the user sends a profile image and self-introduction voice data from the terminal to the server. The server receives the user's profile image and self-introduction voice data and stores them in a database for each user.

[0676] Input: User's profile image (e.g., profile.jpg) and self-introduction audio data (e.g., intro.mp3)

[0677] Output: User profile image and self-introduction audio data stored in the server database

[0678] Specific operation: The user selects a profile picture and voice data on the device and clicks the "Send" button. The server associates the received data with the user ID and stores it in the database.

[0679] Step 2:

[0680] The sending device captures the user's camera video and microphone audio in real time during the online conference, so that the user's current video and audio are input to the device.

[0681] Input: Real-time captured video and audio data

[0682] Output: Video frames and audio samples captured in real time

[0683] Specific operation: The camera captures the user's image and the microphone records the user's voice. These data are stored in the device's memory.

[0684] Step 3:

[0685] The transmitting terminal extracts facial landmarks (such as the positions of the eyes, nose, and mouth) and gestures from the captured video data, and frequency spectrum characteristics from the audio data.

[0686] Input: Video frames and audio samples captured in real time

[0687] Output: Extracted facial landmarks (e.g., eye position (x1, y1), nose position (x2, y2), mouth position (x3, y3)), frequency spectrum characteristics (e.g., frequency band Hz, amplitude dB)

[0688] How it works: The image processing algorithm analyzes the video data and detects facial features, while the audio analysis algorithm analyzes the frequency characteristics of the audio data.

[0689] Step 4:

[0690] The transmitting terminal uses an emotion engine to recognize emotions from the user's facial expressions and tone of voice, and generates emotion data.

[0691] Input: Video frames and audio samples captured in real time

[0692] Output: Recognized emotion data (e.g., "joy", "anger")

[0693] Specific operation: The emotion recognition engine analyzes facial expressions and voice to identify the user's emotional state, resulting in emotional data.

[0694] Step 5:

[0695] The transmitting terminal compresses the extracted feature data and emotion data and transmits them as transmission packets to the server.

[0696] Input: extracted feature data and emotion data

[0697] Output: Compressed outgoing packets

[0698] Specific operation: The data compression algorithm compresses the feature data and emotion data to generate a transmission packet, which the device then sends to the server.

[0699] Step 6:

[0700] The server analyzes the transmitted feature data and emotion data and transmits them to the receiving terminal in an appropriate format.

[0701] Input: Outgoing packets

[0702] Output: Data reformatted to the appropriate format

[0703] Specific operation: The server decompresses and analyzes the transmitted packets, reformats the analysis results into a format usable by the receiving terminal, and transmits them.

[0704] Step 7:

[0705] The receiving terminal inputs the feature data and emotion data sent from the server into the generative AI model and generates video and audio in real time.

[0706] Input: Feature data and emotion data sent from the server

[0707] Output: Generated real-time video and audio

[0708] How it works: A prompt is input into the generative AI model, which generates video and audio in real time, with the recognized emotion reflected in the generated video and audio.

[0709] The above processing steps enable efficient data transmission and reception and natural communication in online conferences and video streaming.

[0710] (Application example 2)

[0711] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0712] Online conferences and video streaming services that utilize real-time video and audio data require the transmission and reception of large amounts of data, resulting in increased communication costs and network latency. Furthermore, it is difficult to accurately convey users' emotions, which can impede natural communication. This creates a demand for more efficient and natural communication methods.

[0713] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0714] In this invention, the server includes a sending terminal that extracts feature data from real-time video and audio data, a means for compressing the feature data and emotion data into transmission packets, a means for transmitting the transmission packets to a receiving terminal, and a means for the receiving terminal to generate real-time video and audio using a generation AI based on the feature data and emotion data. This makes it possible to generate natural video and audio that reflects the user's emotions while significantly reducing the amount of data sent and received.

[0715] The "transmitting terminal" is a device where a user generates video and audio data in real time and extracts this as feature data and emotion data.

[0716] "Feature data" is analytical information such as facial landmarks and audio frequency characteristics extracted from real-time video and audio data.

[0717] "Emotion data" is information obtained as a result of analyzing the user's emotional state based on the extracted feature data.

[0718] A "transmission packet" is a data packet containing compressed feature data and emotion data, which is transmitted over a network.

[0719] The "recipient terminal" is a device that receives the transmitted packets and runs the generation AI to generate video and audio in real time.

[0720] "Generative AI" is an artificial intelligence technology that generates video and audio in real time based on received feature data and emotional data.

[0721] A "specialized algorithm" is a program that uses specific processing techniques to extract feature data and emotion data from video and audio data.

[0722] A "machine learning model" is a trained model that, as part of generative AI, generates video and audio in real time based on incoming data.

[0723] MODE FOR CARRYING OUT THE INVENTION

[0724] An object of the present invention is to provide a system that improves the efficiency of data transmission and reception in online conferences and live streaming services, and realizes natural communication that reflects the user's emotions. The following describes in detail an embodiment of the present invention.

[0725] Program processing overview

[0726] The system of the present invention consists of three main components: a sending terminal, a server, and a receiving terminal.

[0727] Sending device

[0728] The transmitting device captures real-time video and audio and extracts feature and emotion data from it. Specifically, it acquires video and audio using hardware such as a camera and microphone, and analyzes facial landmarks and audio frequency characteristics using a dedicated algorithm. Based on the results of this analysis, the emotion engine recognizes the user's emotions. The extracted feature and emotion data are compressed and sent to the server as a transmission packet.

[0729] server

[0730] The server manages and analyzes the feature data and emotion data received from the sending device and sends it to the receiving device. To ensure efficient data transmission and reception, the server converts the feature data and emotion data into an appropriate format and prepares it in a form that is easy for the generation AI to process before sending it.

[0731] Recipient's device

[0732] The receiver's device runs a generation AI based on the data received from the server to generate real-time video and audio. The generation AI uses a machine learning model to reproduce the feature data and emotion data compressed for transmission in real time, generating natural-looking video and audio that also reflects the user's emotions.

[0733] Hardware and Software Used

[0734] Sending device: camera, microphone, specialized algorithms (e.g., facial landmark detection algorithms)

[0735] Emotion engine: Software for recognizing emotions from audio and video (e.g., EmotionEngine)

[0736] Server: Server software that manages and analyzes data transmission and reception

[0737] Recipient device: Generative AI model (e.g., a generative AI using a specific machine learning model)

[0738] Data processing and calculation

[0739] Processing on the sending device: Facial landmarks and audio frequency characteristics are extracted from video and audio captured in real time, and emotions are recognized using an emotion engine.

[0740] Data compression and transmission: The extracted feature data and emotion data are compressed and sent to the server as a transmission packet.

[0741] Processing on the server: The received data is analyzed, converted into a format that is easy for the generating AI to process, and sent to the recipient's device.

[0742] Processing on the recipient's device: Using generative AI, video and audio are generated in real time based on the transmitted feature data and emotion data.

[0743] Specific examples

[0744] When a live streamer broadcasts live on their smartphone, they register a profile picture and audio sample in advance. During the live broadcast, the sending device captures video and audio in real time, extracts feature data and emotional data, and sends it to a server. The receiving device runs a generative AI based on the received data to generate video and audio in real time. This method makes it possible to provide viewers with high-quality video and audio with low latency.

[0745] Example prompt sentence:

[0746] "Generate video and audio in real time based on the user's profile picture and voice sample. Recognize the user's emotions from facial expressions and tone of voice."

[0747] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0748] Step 1:

[0749] The sending device captures video and audio in real time from the user's camera and microphone. The input is the camera video and microphone audio, and the output is the captured video frames and audio samples of these data. Specifically, it performs a process to acquire the data stream from the camera and microphone.

[0750] Step 2:

[0751] The transmitting device detects facial landmarks (such as the eyes, nose, and mouth) from the captured video and extracts frequency spectrum characteristics from the audio. The input is video frames and audio samples, and the output is facial landmark data and audio characteristic data. Specifically, it runs a facial landmark detection algorithm and an audio spectrum analysis algorithm.

[0752] Step 3:

[0753] The transmitting device uses an emotion engine to recognize the user's emotion based on the detected facial landmarks and voice characteristic data. The input is facial landmark data and voice characteristic data, and the output is recognized emotion data. Specifically, the emotion engine is executed to analyze the user's emotion from their facial expression and voice tone.

[0754] Step 4:

[0755] The transmitting terminal compresses the extracted feature data and emotion data into transmission packets. The input is facial landmark data, voice characteristic data, and emotion data, and the output is compressed transmission packets. Specifically, the data is efficiently packetized using a data compression algorithm.

[0756] Step 5:

[0757] The sending terminal sends compressed transmission packets to the server. The input is the transmission packet, and the output is the data sent to the server. Specifically, the data is sent to the server using a network protocol.

[0758] Step 6:

[0759] The server decompresses the received transmission packets and analyzes the feature data and emotion data. The input is the received transmission packets, and the output is the decompressed feature data and emotion data. Specifically, the server executes a process to decompress the packets using a data decompression algorithm.

[0760] Step 7:

[0761] The server transmits the feature data and emotion data to the recipient's terminal. The input is the decompressed feature data and emotion data, and the output is the data transmitted to the recipient's terminal. Specifically, the data is transmitted using a network protocol.

[0762] Step 8:

[0763] The receiver's device inputs the received feature data and emotion data into the generative AI to generate real-time video and audio. The input is the received feature data and emotion data, and the output is the generated video and audio. Specifically, the generative AI model is used to process the video and audio based on the feature data and emotion data.

[0764] Step 9:

[0765] The receiver's terminal presents the generated real-time video and audio to the user. The input is the generated video and audio, and the output is the video and audio presented to the user. Specifically, the system performs a process to play back the generated content using a display and speakers.

[0766] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0767] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0768] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0769] [Third embodiment]

[0770] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0771] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0772] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0773] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0774] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0775] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0776] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0777] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0778] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0779] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0780] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0781] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0782] The system of the present invention utilizes generative AI to generate video and audio in real time based on feature data in order to improve the efficiency of data transmission and reception in online conferences and video streaming.

[0783] Program processing overview

[0784] User connection and pre-transmission of data

[0785] The server receives profile images and voice samples of users when they connect to an online conference. During this process, users submit their own images and voice data introducing themselves, which are then stored on the server.

[0786] Extracting and sending feature data

[0787] The transmitting device captures the user's camera video and microphone audio in real time and extracts feature data from them, including facial landmarks (such as the positions of the eyes, nose, and mouth) and frequency spectrum characteristics of the voice.

[0788] The terminal compresses the extracted feature data and transmits it to the server as an efficient transmission packet.

[0789] Video Creation and Distribution

[0790] The server analyzes the feature data received from the sending device and transmits it to the receiving device in an appropriate format, allowing the generating AI to recreate the video and audio in real time.

[0791] The receiving device runs the generative AI using the feature data sent from the server to generate video and audio in real time, which is then displayed to the receiving user, providing an experience similar to that of a regular online meeting.

[0792] Specific examples

[0793] 1. Connection and Pre-Data Transmission

[0794] User A connects to the online conference, and the server receives the profile picture and self-introduction voice of User A. User A's terminal transmits this data to the server.

[0795] 2. Extracting and sending feature data

[0796] The transmitting terminal (User A's device) captures User A's video and audio in real time during the conference and extracts feature data such as facial position information and audio frequency characteristics. This feature data is efficiently compressed and sent to the server as a transmission packet.

[0797] 3. Video Creation and Distribution

[0798] The server analyzes the received feature data and sends it in an appropriate format to the receiving device (User B's device). The receiving device inputs the transmitted feature data into the generation AI, which generates real-time video and audio of User A. User B then watches the generated video and audio, enjoying a natural online meeting experience.

[0799] In this way, the system of the present invention can provide real-time video and audio while significantly reducing the amount of data transmitted and received and easing the strain on network resources, resulting in cost reduction and a reduction in the environmental impact.

[0800] The processing flow will be explained below.

[0801] Step 1:

[0802] A user connects to an online meeting.

[0803] The user launches the online meeting application and connects to the server.

[0804] Step 2:

[0805] The server receives the advance data.

[0806] The server receives a profile image and voice sample from the user, which includes a photo of the user's face and an audio introduction.

[0807] Step 3:

[0808] The user submits the advance data.

[0809] Users send their profile picture and voice sample from their device to the server, and this data is sent when the meeting is initially connected.

[0810] Step 4:

[0811] The device captures video and audio.

[0812] The sending device captures the user's camera video and microphone audio in real time, allowing conversations and actions to be recorded.

[0813] Step 5:

[0814] The device extracts the feature data.

[0815] The device extracts facial landmarks (such as the positions of the eyes, nose, and mouth) and gestures from the captured video, and also extracts frequency spectrum characteristics from the audio data.

[0816] Step 6:

[0817] The terminal compresses the feature data.

[0818] The terminal efficiently compresses the extracted feature data to generate transmission packets, which significantly reduces the amount of data.

[0819] Step 7:

[0820] The terminal transmits the feature data to the server.

[0821] The terminal transmits a transmission packet to the server in real time, and the transmission packet includes compressed feature data.

[0822] Step 8:

[0823] The server analyzes the feature data.

[0824] The server parses the received feature data and decodes it into the appropriate format, making the data usable by the receiving device.

[0825] Step 9:

[0826] The server transmits the data to the receiving terminal.

[0827] The server transmits the analyzed feature data to the receiving terminal, which then generates the data in real time.

[0828] Step 10:

[0829] The receiving device launches the generation AI.

[0830] The receiving device activates the generation AI based on the transmitted feature data and inputs the data, which starts the generation process.

[0831] Step 11:

[0832] Generative AI generates real-time video and audio.

[0833] The generative AI uses the feature data to generate real-time video and audio that replicates the movements and mouth movements of the sending user.

[0834] Step 12:

[0835] The receiving terminal displays the generated video and audio to the user.

[0836] The receiving terminal displays the generated video and audio to the user in real time, and the user can see and hear the video and audio of other participants just like in a regular online conference.

[0837] This system significantly reduces the amount of data sent and received, easing the load on network resources and resulting in cost savings and a reduced environmental impact.

[0838] Example 1

[0839] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0840] In online conferences and video streaming, sending and receiving real-time video and audio data requires a large amount of network resources, resulting in communication delays and quality degradation. Furthermore, existing technologies lack efficient methods for generating high-quality video and audio in real time. This results in poor user experience and increases costs and environmental impact.

[0841] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0842] In this invention, the server includes a means for a user to connect to an online conference and transmit a profile image and a voice sample, a transmitting terminal including means for extracting feature data including facial landmarks and voice frequency spectrum characteristics from real-time video and audio data, a means for compressing the feature data into a transmission packet, a means for transmitting the transmission packet to the server and for the server to convert the feature data into an appropriate format, a means for the server to transmit the converted data to a receiving terminal, and a means for the receiving terminal to generate real-time video and audio using a generative AI model based on the feature data. This significantly reduces the amount of data transmitted and received, enabling real-time generation of high-quality video and audio while minimizing communication delays.

[0843] "Users" are typical participants who use online meetings and video streaming services.

[0844] An "online conference" refers to a conference or meeting that takes place over the Internet in real time, sharing audio and video.

[0845] A "profile image" is a still image that a user uses to indicate their personal information or identity.

[0846] A "voice sample" is a short piece of voice data that a user provides to the server, used for self-introduction and identification.

[0847] The "transmitting terminal" is a device used by a user (for example, a smartphone or a personal computer) that captures video and audio using a camera and microphone.

[0848] "Feature data" are specific digital indices or characteristics (e.g., facial landmarks or frequency spectrum characteristics of voice) extracted from captured video and audio data.

[0849] "Landmarks" are characteristic points that indicate the position information of each part of the face (eyes, nose, mouth, etc.).

[0850] The "frequency spectrum characteristics" are data that indicate the strength of each frequency component in an audio signal.

[0851] A "server" is a central system that manages and processes data sent and received from sending and receiving terminals.

[0852] A "transmission packet" is a small unit of compressed data containing feature data that is transmitted over a network.

[0853] The "receiving terminal" is a device that receives the feature data sent from the server and reproduces real-time video and audio using a generation AI.

[0854] A "generative AI model" is an artificial intelligence technology that uses machine learning algorithms to generate video and audio from feature data.

[0855] "Compression" is a technique for reducing data size and is used to improve transmission efficiency.

[0856] "Real-time" refers to instantaneous processing and communication with little or no delay.

[0857] MODE FOR CARRYING OUT THE INVENTION

[0858] The system of the present invention utilizes a generative AI model to generate video and audio in real time based on feature data in order to improve the efficiency of data transmission and reception in online meetings and video streaming. A specific embodiment of this system is described below.

[0859] Hardware and software used

[0860] Hardware: A computer device with a compatible camera, microphone, and internet connection

[0861] Software: Generative AI models (e.g., GPT-3, DALL-E), libraries for video and audio capture (e.g., OpenCV, PyAudio)

[0862] A natural language description of the program

[0863] User connection and pre-transmission of data

[0864] A user opens an online conference app on their computer and connects to the server. The user then sends a profile picture and a self-introduction voice sample. During this process, the user follows the application's instructions to select a profile picture they have taken in advance, record a self-introduction voice using a microphone, and send it to the server. The server then stores the received data.

[0865] Real-time extraction and transmission of feature data

[0866] The sending device (e.g., User A's device) uses a camera and microphone to capture User A's video and audio in real time. Specifically, it uses a computer vision library such as OpenCV to detect facial landmarks and an audio analysis library such as PyAudio to analyze the frequency spectrum characteristics of the audio. These feature data are efficiently compressed and sent to the server as transmission packets.

[0867] Feature data analysis and resubmission

[0868] The server analyzes the feature data received from the sending device and converts the data into the appropriate format. This involves using generative AI models to calculate additional data needed to generate the video and audio. For example, it generates data to complete facial details and audio. The server then transmits the converted data to the receiving device.

[0869] Real-time video and audio generation

[0870] The receiving device (e.g., User B's device) receives the feature data and complementary data sent from the server, and the internal generative AI model is activated. For example, DALL-E generates a video reproducing User A's face based on the facial landmark data and complementary data, and simultaneously generates real-time audio based on the audio sample. The generated video and audio are output to the screen and speakers, just like a regular video conferencing app.

[0871] Specific examples

[0872] 1. Connection and Pre-Data Transmission

[0873] User A connects to the online conference, and the server receives the profile picture and self-introduction voice of User A. User A's terminal transmits this data to the server.

[0874] 2. Real-time extraction and transmission of feature data

[0875] The transmitting terminal (User A's device) captures User A's video and audio in real time during the conference and extracts feature data such as facial position information and audio frequency characteristics. This feature data is efficiently compressed and sent to the server as a transmission packet.

[0876] 3. Analysis and retransmission of feature data

[0877] The server analyzes the received feature data and sends it in an appropriate format to the receiving device of User B. User B's device inputs the transmitted feature data into the generation AI, which generates real-time video and audio of User A.

[0878] Prompt Sentence Examples

[0879] Data sent from User A's device to the server: "Send profile picture and voice sample"

[0880] The process in which the server sends feature data to User B's device: "Generate video and audio based on facial landmarks and audio spectrum characteristics."

[0881] This system significantly reduces the amount of data transmitted and received, alleviating pressure on network resources while providing high-quality real-time video and audio, resulting in cost savings and a lighter environmental impact.

[0882] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0883] Step 1:

[0884] User connection and pre-transmission of data

[0885] A user opens an online conference app on their computer and connects to the server. The user provides a profile image and a self-introduction voice sample as input, which is then sent to the server. The input image and voice data are stored on the server. For example, the user may upload a face photo taken with a camera or a voice recording with a microphone. This process is performed to collect initial data that will serve as the basis for subsequent feature data generation.

[0886] Step 2:

[0887] Real-time extraction and transmission of feature data

[0888] The sending terminal (user's device) uses a camera and microphone to capture the user's video and audio in real time. This becomes the input data. Specifically, a computer vision library such as OpenCV is used to detect facial landmarks (the positions of the eyes, nose, mouth, etc.), and an audio analysis library such as PyAudio is used to analyze the frequency spectrum characteristics of the audio. This feature data is compressed and sent to the server as a transmission packet. For example, this includes operations to obtain the position information of each part of the face from the camera video and calculate the strength of specific frequency components from the audio signal. The output is compressed feature data.

[0889] Step 3:

[0890] Feature data analysis and resubmission

[0891] The server receives feature data from the sending device as input. It analyzes this data and uses a generative AI model (e.g., GPT-3 or DALL-E) to calculate the necessary completion data for video and audio generation. Specifically, it generates completion data for facial details and audio. This data is then compressed again and sent to the receiving device in an appropriate format. For example, this may involve the generative AI completing fine details of the user's face to generate high-quality video data. The output is a packet containing the completion data and the original feature data.

[0892] Step 4:

[0893] Real-time video and audio generation

[0894] The receiving device (User B's device) receives the feature data and complementary data from the server as input. An internal generative AI model is activated and generates video and audio in real time based on this data. Specifically, DALL-E generates video recreating User A's face based on facial landmark data, and generates real-time audio based on audio samples. The output is the generated video and audio, which are displayed on User B's screen and played through the speakers.

[0895] In this way, the system of the present invention can provide high-quality video and audio in real time while significantly reducing the amount of data sent and received and easing the strain on network resources.

[0896] (Application example 1)

[0897] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0898] Conventional online conference systems and video streaming systems require large amounts of data for transmission and reception, resulting in problems such as strain on network resources and communication delays. Furthermore, real-time natural dialogue between customers and store staff in virtual stores requires the transmission and reception of high-quality video and audio. However, this data transmission depends on bandwidth and communication quality, which can potentially impair the customer experience. A system that can solve this issue and efficiently provide real-time video and audio is needed.

[0899] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0900] In this invention, the server includes a means for receiving and storing customer profile images and voice samples, a means for capturing video and audio of the virtual store clerk in real time, extracting and compressing feature data, and transmitting the data to the server, and a means for the receiving terminal to play back video and audio of the virtual store clerk in real time using artificial intelligence for generation, thereby enabling efficient use of network resources and providing high-quality real-time video and audio.

[0901] The "transmitting terminal" is a terminal that captures video and audio data and extracts feature data in real time.

[0902] "Feature data" is data extracted from video and audio data, such as facial landmarks and audio frequency spectrum characteristics.

[0903] The "means for compressing and forming transmission packets" refers to a means for efficiently compressing the extracted feature data and converting it into a packet format for transmission.

[0904] The "recipient's terminal" is a terminal that receives the feature data sent from the sender and generates video and audio in real time using artificial intelligence.

[0905] "Generative AI" is an AI technology that generates video and audio in real time based on feature data.

[0906] A "virtual store" is a store that exists in a virtual space and that users can access online to purchase goods and services.

[0907] A "profile picture" is a static image that a user provides when connecting to an online service.

[0908] "Audio sample" is audio data provided by a user when connecting to an online service.

[0909] A "virtual store clerk" is a virtual character that serves customers in a virtual store.

[0910] "Means for playing in real time" refers to means for instantly generating video and audio using artificial intelligence based on received feature data, and displaying and playing them to the user.

[0911] The system for implementing this invention includes a real-time customer service system in a virtual store. The system is composed of the following main components:

[0912] Server-side processing

[0913] When a customer visits the virtual store, the server receives and stores the customer's profile image and voice sample (e.g., self-introduction). The profile data is used for customer identification and data matching. Next, the server receives real-time video and audio data sent from the store clerk's terminal and extracts feature data from that data. Specifically, this includes facial landmarks and frequency spectrum characteristics of the voice. The extracted feature data is efficiently compressed and sent to the recipient's terminal.

[0914] Terminal side processing

[0915] The recipient's device (customer's device) receives the feature data sent from the server. Based on the received feature data, a generative AI model is used to generate video and audio of a virtual store clerk in real time. The generated video and audio are then played back to the customer in real time, allowing the customer to enjoy a natural communication experience, as if they were in a real store.

[0916] Hardware and software used

[0917] Server: A server with powerful CPU and GPU. Uses generative artificial intelligence models (e.g., OpenAI's GPT series).

[0918] Client device: Smartphone or head-mounted display (HMD). The OpenCV library is used to receive feature data and play back the generated video and audio.

[0919] Specific examples

[0920] Example prompt:

[0921] "When customers visit a virtual store, they communicate with a virtual associate in real time based on their visit."

[0922] This invention solves the data transmission and reception issues of conventional online conference systems and video streaming systems. It also enables real-time transmission of high-quality video and audio in virtual stores, improving the customer experience and enabling efficient use of network resources.

[0923] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0924] Step 1:

[0925] When a user visits the virtual store, the server receives and stores the user's profile picture and voice sample (e.g., self-introduction), which are used to identify the user and provide customized services.

[0926] Input: User's profile picture and voice sample.

[0927] Output: User profile data stored on the server.

[0928] How it works: The server stores user-submitted profile images and voice samples in a database.

[0929] Step 2:

[0930] The virtual store assistant uses a dedicated device to capture his or her own video and audio in real time, and the transmitting device extracts feature data such as facial landmarks and frequency spectrum characteristics from this data, then efficiently compresses and converts it into transmission packets.

[0931] Input: Real-time video and audio data.

[0932] Output: Outgoing packets containing compressed feature data.

[0933] How it works: The sending device analyzes the video and audio data to extract facial landmarks and audio frequency spectrum characteristics, then compresses and converts them into packets for transmission.

[0934] Step 3:

[0935] The server receives the transmitted packets from the sending terminal, analyzes them, and routes them appropriately to each receiving terminal.

[0936] Input: Transmission packets containing compressed feature data.

[0937] Output: Parsed feature data.

[0938] How it works: The server unpacks the received transmission packet and analyzes the feature data against the profile data stored in its internal database.

[0939] Step 4:

[0940] The receiver's device receives the feature data sent from the server and uses the generative AI model to generate real-time video and audio, which is then instantly played back to the user (customer).

[0941] Input: Feature data sent from the server.

[0942] Output: Real-time generated video and audio.

[0943] How it works: The recipient's device inputs the received feature data into a generative AI model, which then generates and plays back video and audio of a virtual store clerk based on this data.

[0944] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0945] The system of the present invention utilizes generative AI and an emotion engine to generate video and audio in real time and recognize emotions based on feature data in order to improve the efficiency of data transmission and reception in online meetings and video streaming.

[0946] Program processing overview

[0947] User connection and pre-transmission of data

[0948] The server receives profile images and voice samples of users when they connect to an online conference. During this process, users submit their own images and voice data introducing themselves, which are then stored on the server.

[0949] Extracting and sending feature data

[0950] The transmitting device captures the user's camera video and microphone audio in real time and extracts feature data from them, specifically facial landmarks (such as the positions of the eyes, nose, and mouth), gestures, and frequency spectrum characteristics from audio data.

[0951] The device recognizes the user's emotions based on the captured video and audio data using an emotion engine that analyzes emotions from facial expressions and voice tone.

[0952] The terminal compresses the extracted feature data and the recognized emotion data and transmits them to the server as efficient transmission packets.

[0953] Video Creation and Distribution

[0954] The server analyzes the feature and emotion data received from the sending device and transmits it to the receiving device in an appropriate format, allowing the generating AI to reproduce the video and audio in real time.

[0955] The receiving device uses the feature data and emotion data sent from the server to run the generative AI and generate video and audio in real time. The generated video and audio also reflect the recognized emotions, realizing natural communication.

[0956] Specific examples

[0957] 1. Connection and Pre-Data Transmission

[0958] User A connects to the online conference, and the server receives the profile picture and self-introduction voice of User A. User A's terminal transmits this data to the server.

[0959] 2. Extracting and sending feature data

[0960] The transmitting terminal (User A's device) captures the video and audio of User A in real time during the conference. From this data, feature data such as facial position information and audio frequency characteristics are extracted.

[0961] Furthermore, the emotion engine installed in User A's device recognizes User A's emotions from facial expressions and voice tone, analyzing emotions such as smiling or anger, for example.

[0962] The extracted feature data and emotion data are compressed and transmitted to the server as transmission packets.

[0963] 3. Video Creation and Distribution

[0964] The server analyzes the received feature data and emotion data and sends them in an appropriate format to the receiving device (User B's device). The receiving device inputs the transmitted data into the generation AI, which generates real-time video and audio of User A.

[0965] The generated video and audio also reflect User A's emotions, allowing User B to have a natural online meeting experience based on User A's facial expressions and tone.

[0966] In this way, the system of the present invention can provide natural communication through emotion recognition while significantly reducing the amount of data sent and received and easing the strain on network resources, resulting in cost reduction and a reduction in the environmental impact.

[0967] The processing flow will be explained below.

[0968] Step 1:

[0969] A user connects to an online meeting.

[0970] The user launches an online conference application and connects to the server by entering a URL or ID to connect to a specific conference room.

[0971] Step 2:

[0972] The server receives the advance data.

[0973] Upon initial connection, the server receives a profile image and voice sample from the user, including a photo of the user's face and a voice clip introducing themselves.

[0974] Step 3:

[0975] The user submits the advance data.

[0976] The user device sends a facial photo taken with a camera and a voice sample recorded with a microphone to the server, and this pre-data is used for user identification and initial setup.

[0977] Step 4:

[0978] The device captures video and audio.

[0979] The sending device captures the user's camera video and microphone audio in real time, recording the video and audio data frame by frame.

[0980] Step 5:

[0981] The device extracts the feature data.

[0982] The device detects facial landmarks (such as the eyes, nose, and mouth) and gestures from the captured video. Frequency spectrum characteristics are also extracted from the audio data. This allows the device to accurately understand the user's movements and speech while reducing the amount of data.

[0983] Step 6:

[0984] The device runs the emotion engine.

[0985] The device's built-in emotion engine recognizes the user's emotions from captured video and audio data, using a dedicated algorithm to identify emotions such as joy, anger, sadness, and happiness from facial expressions and voice tone.

[0986] Step 7:

[0987] The device compresses the feature data and emotion data.

[0988] The device efficiently compresses the extracted feature data and emotion data and sends them as packets, which are small in size and minimize network load.

[0989] Step 8:

[0990] The terminal transmits a transmission packet to the server.

[0991] The device sends packets containing compressed feature data and emotion data to the server in real time, a process that is fast and causes almost no delay.

[0992] Step 9:

[0993] The server analyzes the feature data and emotion data.

[0994] The server analyzes the received transmission packets and decodes the feature data and emotion data, converting them into a format that can be used by the receiving terminal.

[0995] Step 10:

[0996] The server transmits the data to the receiving terminal.

[0997] The server then transmits the analyzed feature data and emotion data to the receiving device, which receives this data and uses it as input for the generation AI.

[0998] Step 11:

[0999] The receiving device launches the generation AI.

[1000] The receiving device inputs the transmitted feature data and emotion data into the generation AI, initiating the real-time video and audio generation process.

[1001] Step 12:

[1002] Generative AI generates real-time video and audio.

[1003] The generative AI generates real-time video and audio of the user from feature data, and also reflects the user's facial expressions and vocal intonation based on emotional data.

[1004] Step 13:

[1005] The receiving terminal displays the generated video and audio to the user.

[1006] The receiving terminal displays the generated video and audio to the user in real time, allowing the user to view natural video and audio that reflects the emotions of the sending user.

[1007] This system significantly reduces the amount of data sent and received, alleviating the burden on network resources while providing natural communication through emotion recognition, resulting in cost savings and a lighter environmental impact.

[1008] Example 2

[1009] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1010] In online meetings and video streaming, it is important to generate high-quality video and audio in real time to improve the efficiency of data transmission and reception. It is also necessary to recognize user emotions and realize natural communication.

[1011] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a sending terminal including means for extracting feature data from real-time video and audio data, means for recognizing a user's emotion based on the feature data, means for compressing the feature data and emotion data into transmission packets, means for transmitting the transmission packets to a receiving terminal, and means for the receiving terminal to generate real-time video and audio using a generation AI based on the feature data and emotion data. This enables efficient data transmission and reception and natural communication.

[1012] A "transmitting terminal" is a device that captures video and audio data in real time, extracts feature data from the data, and transmits the data to a server.

[1013] "Feature data" refers to information such as facial landmarks and frequency spectrum characteristics extracted from the user's video and audio data.

[1014] "Emotion data" is information about emotions analyzed from the user's video and audio by an emotion recognition engine.

[1015] A "transmission packet" is a data packet that contains compressed feature data and emotion data.

[1016] The "recipient's terminal" is a device that receives the transmission packets sent from the server and generates video and audio in real time using generation AI.

[1017] "Generative AI" is an artificial intelligence model that generates video and audio in real time based on feature data and emotional data.

[1018] An "emotion recognition engine" is software that analyzes emotions from a user's video and audio data and outputs them as emotional data.

[1019] A "machine learning model" is a model that uses algorithms to learn patterns and regularities from data and make inferences and predictions.

[1020] The system of the present invention improves the efficiency of data transmission and reception in online meetings and video streaming, and utilizes generative AI and an emotion engine to generate video and audio in real time and recognize emotions based on feature data.

[1021] User connection and pre-transmission of data

[1022] When a user connects to an online conference, the user sends their profile picture and self-introduction audio data from their device to the server, which then receives and stores this data for each user.

[1023] Extracting and sending feature data

[1024] The sending device captures the user's camera video and microphone audio in real time during an online conference. From this captured data, the device extracts frequency spectrum characteristics from facial landmarks (such as the positions of the eyes, nose, and mouth), gestures, and voice data. Furthermore, it uses an emotion engine to recognize the user's emotions and generate emotion data.

[1025] The extracted feature data and emotion data are compressed and sent to the server as transmission packets. This compression and transmission process is important for enabling efficient communication.

[1026] Video Creation and Distribution

[1027] The server analyzes the transmitted feature data and emotion data and transmits it to the recipient's device in the appropriate format. This data is formatted so that the generative AI can reproduce the video and audio in real time. The recipient's device inputs the data transmitted from the server into the generative AI model, generating the video and audio in real time. The generated video and audio reflect the user's emotions, enabling natural communication.

[1028] Specific examples

[1029] 1. Connection and Pre-Data Transmission

[1030] When user A connects to an online conference, the terminal sends a profile image (profile.jpg) and self-introduction audio data (intro.mp3) to the server.

[1031] The server receives this data, associates it with User A's ID, and stores it in a database.

[1032] 2. Extracting and sending feature data

[1033] The transmitting device captures the video and audio of User A in real time. From this captured data, facial landmarks (eyes (x1, y1), nose (x2, y2), mouth (x3, y3)) and audio frequency characteristics (frequency band Hz, amplitude dB) are extracted.

[1034] The emotion engine recognizes emotions from user A's facial expressions and tone of voice, and generates emotion data for "joy," for example.

[1035] The extracted feature data and emotion data are compressed and transmitted to the server as transmission packets.

[1036] 3. Video Creation and Distribution

[1037] The server analyzes the received feature data and emotion data and transmits them to the receiving terminal (user B's terminal) in an appropriate format.

[1038] The receiving device inputs the transmitted feature data and emotion data into the generative AI model, generating video and audio of User A in real time. Because the generated video and audio also reflect User A's emotions, User B can enjoy a natural online meeting experience based on User A's facial expressions and tone.

[1039] Prompt Sentence Examples

[1040] "Analyze the camera video and microphone audio of User A, and generate video and audio in real time based on the following feature data. The feature data is as follows:

[1041] Facial landmarks: eyes(x,y), nose(x,y), mouth(x,y)

[1042] Audio frequency spectrum: frequency band (Hz), amplitude (dB)

[1043] Emotion data: smile, anger

[1044] The generated video and audio should reflect the perceived emotions of User A.

[1045] In this way, the system of the present invention can provide efficient data transmission and reception and natural communication, thereby reducing costs and environmental impact.

[1046] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1047] Step 1:

[1048] When a user connects to an online conference, the user sends a profile image and self-introduction voice data from the terminal to the server. The server receives the user's profile image and self-introduction voice data and stores them in a database for each user.

[1049] Input: User's profile image (e.g., profile.jpg) and self-introduction audio data (e.g., intro.mp3)

[1050] Output: User profile image and self-introduction audio data stored in the server database

[1051] Specific operation: The user selects a profile picture and voice data on the device and clicks the "Send" button. The server associates the received data with the user ID and stores it in the database.

[1052] Step 2:

[1053] The sending device captures the user's camera video and microphone audio in real time during the online conference, so that the user's current video and audio are input to the device.

[1054] Input: Real-time captured video and audio data

[1055] Output: Video frames and audio samples captured in real time

[1056] Specific operation: The camera captures the user's image and the microphone records the user's voice. These data are stored in the device's memory.

[1057] Step 3:

[1058] The transmitting terminal extracts facial landmarks (such as the positions of the eyes, nose, and mouth) and gestures from the captured video data, and frequency spectrum characteristics from the audio data.

[1059] Input: Video frames and audio samples captured in real time

[1060] Output: Extracted facial landmarks (e.g., eye position (x1, y1), nose position (x2, y2), mouth position (x3, y3)), frequency spectrum characteristics (e.g., frequency band Hz, amplitude dB)

[1061] How it works: The image processing algorithm analyzes the video data and detects facial features, while the audio analysis algorithm analyzes the frequency characteristics of the audio data.

[1062] Step 4:

[1063] The transmitting terminal uses an emotion engine to recognize emotions from the user's facial expressions and tone of voice, and generates emotion data.

[1064] Input: Video frames and audio samples captured in real time

[1065] Output: Recognized emotion data (e.g., "joy", "anger")

[1066] Specific operation: The emotion recognition engine analyzes facial expressions and voice to identify the user's emotional state, resulting in emotional data.

[1067] Step 5:

[1068] The transmitting terminal compresses the extracted feature data and emotion data and transmits them as transmission packets to the server.

[1069] Input: extracted feature data and emotion data

[1070] Output: Compressed outgoing packets

[1071] Specific operation: The data compression algorithm compresses the feature data and emotion data to generate a transmission packet, which the device then sends to the server.

[1072] Step 6:

[1073] The server analyzes the transmitted feature data and emotion data and transmits them to the receiving terminal in an appropriate format.

[1074] Input: Outgoing packets

[1075] Output: Data reformatted to the appropriate format

[1076] Specific operation: The server decompresses and analyzes the transmitted packets, reformats the analysis results into a format usable by the receiving terminal, and transmits them.

[1077] Step 7:

[1078] The receiving terminal inputs the feature data and emotion data sent from the server into the generative AI model and generates video and audio in real time.

[1079] Input: Feature data and emotion data sent from the server

[1080] Output: Generated real-time video and audio

[1081] How it works: A prompt is input into the generative AI model, which generates video and audio in real time, with the recognized emotion reflected in the generated video and audio.

[1082] The above processing steps enable efficient data transmission and reception and natural communication in online conferences and video streaming.

[1083] (Application example 2)

[1084] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1085] Online conferences and video streaming services that utilize real-time video and audio data require the transmission and reception of large amounts of data, resulting in increased communication costs and network latency. Furthermore, it is difficult to accurately convey users' emotions, which can impede natural communication. This creates a demand for more efficient and natural communication methods.

[1086] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1087] In this invention, the server includes a sending terminal that extracts feature data from real-time video and audio data, a means for compressing the feature data and emotion data into transmission packets, a means for transmitting the transmission packets to a receiving terminal, and a means for the receiving terminal to generate real-time video and audio using a generation AI based on the feature data and emotion data. This makes it possible to generate natural video and audio that reflects the user's emotions while significantly reducing the amount of data sent and received.

[1088] The "transmitting terminal" is a device where a user generates video and audio data in real time and extracts this as feature data and emotion data.

[1089] "Feature data" is analytical information such as facial landmarks and audio frequency characteristics extracted from real-time video and audio data.

[1090] "Emotion data" is information obtained as a result of analyzing the user's emotional state based on the extracted feature data.

[1091] A "transmission packet" is a data packet containing compressed feature data and emotion data, which is transmitted over a network.

[1092] The "recipient terminal" is a device that receives the transmitted packets and runs the generation AI to generate video and audio in real time.

[1093] "Generative AI" is an artificial intelligence technology that generates video and audio in real time based on received feature data and emotional data.

[1094] A "specialized algorithm" is a program that uses specific processing techniques to extract feature data and emotion data from video and audio data.

[1095] A "machine learning model" is a trained model that, as part of generative AI, generates video and audio in real time based on incoming data.

[1096] MODE FOR CARRYING OUT THE INVENTION

[1097] An object of the present invention is to provide a system that improves the efficiency of data transmission and reception in online conferences and live streaming services, and realizes natural communication that reflects the user's emotions. The following describes in detail an embodiment of the present invention.

[1098] Program processing overview

[1099] The system of the present invention consists of three main components: a sending terminal, a server, and a receiving terminal.

[1100] Sending device

[1101] The transmitting device captures real-time video and audio and extracts feature and emotion data from it. Specifically, it acquires video and audio using hardware such as a camera and microphone, and analyzes facial landmarks and audio frequency characteristics using a dedicated algorithm. Based on the results of this analysis, the emotion engine recognizes the user's emotions. The extracted feature and emotion data are compressed and sent to the server as a transmission packet.

[1102] server

[1103] The server manages and analyzes the feature data and emotion data received from the sending device and sends it to the receiving device. To ensure efficient data transmission and reception, the server converts the feature data and emotion data into an appropriate format and prepares it in a form that is easy for the generation AI to process before sending it.

[1104] Recipient's device

[1105] The receiver's device runs a generation AI based on the data received from the server to generate real-time video and audio. The generation AI uses a machine learning model to reproduce the feature data and emotion data compressed for transmission in real time, generating natural-looking video and audio that also reflects the user's emotions.

[1106] Hardware and Software Used

[1107] Sending device: camera, microphone, specialized algorithms (e.g., facial landmark detection algorithms)

[1108] Emotion engine: Software for recognizing emotions from audio and video (e.g., EmotionEngine)

[1109] Server: Server software that manages and analyzes data transmission and reception

[1110] Recipient device: Generative AI model (e.g., a generative AI using a specific machine learning model)

[1111] Data processing and calculation

[1112] Processing on the sending device: Facial landmarks and audio frequency characteristics are extracted from video and audio captured in real time, and emotions are recognized using an emotion engine.

[1113] Data compression and transmission: The extracted feature data and emotion data are compressed and sent to the server as a transmission packet.

[1114] Processing on the server: The received data is analyzed, converted into a format that is easy for the generating AI to process, and sent to the recipient's device.

[1115] Processing on the recipient's device: Using generative AI, video and audio are generated in real time based on the transmitted feature data and emotion data.

[1116] Specific examples

[1117] When a live streamer broadcasts live on their smartphone, they register a profile picture and audio sample in advance. During the live broadcast, the sending device captures video and audio in real time, extracts feature data and emotional data, and sends it to a server. The receiving device runs a generative AI based on the received data to generate video and audio in real time. This method makes it possible to provide viewers with high-quality video and audio with low latency.

[1118] Example prompt sentence:

[1119] "Generate video and audio in real time based on the user's profile picture and voice sample. Recognize the user's emotions from facial expressions and tone of voice."

[1120] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1121] Step 1:

[1122] The sending device captures video and audio in real time from the user's camera and microphone. The input is the camera video and microphone audio, and the output is the captured video frames and audio samples of these data. Specifically, it performs a process to acquire the data stream from the camera and microphone.

[1123] Step 2:

[1124] The transmitting device detects facial landmarks (such as the eyes, nose, and mouth) from the captured video and extracts frequency spectrum characteristics from the audio. The input is video frames and audio samples, and the output is facial landmark data and audio characteristic data. Specifically, it runs a facial landmark detection algorithm and an audio spectrum analysis algorithm.

[1125] Step 3:

[1126] The transmitting device uses an emotion engine to recognize the user's emotion based on the detected facial landmarks and voice characteristic data. The input is facial landmark data and voice characteristic data, and the output is recognized emotion data. Specifically, the emotion engine is executed to analyze the user's emotion from their facial expression and voice tone.

[1127] Step 4:

[1128] The transmitting terminal compresses the extracted feature data and emotion data into transmission packets. The input is facial landmark data, voice characteristic data, and emotion data, and the output is compressed transmission packets. Specifically, the data is efficiently packetized using a data compression algorithm.

[1129] Step 5:

[1130] The sending terminal sends compressed transmission packets to the server. The input is the transmission packet, and the output is the data sent to the server. Specifically, the data is sent to the server using a network protocol.

[1131] Step 6:

[1132] The server decompresses the received transmission packets and analyzes the feature data and emotion data. The input is the received transmission packets, and the output is the decompressed feature data and emotion data. Specifically, the server executes a process to decompress the packets using a data decompression algorithm.

[1133] Step 7:

[1134] The server transmits the feature data and emotion data to the recipient's terminal. The input is the decompressed feature data and emotion data, and the output is the data transmitted to the recipient's terminal. Specifically, the data is transmitted using a network protocol.

[1135] Step 8:

[1136] The receiver's device inputs the received feature data and emotion data into the generative AI to generate real-time video and audio. The input is the received feature data and emotion data, and the output is the generated video and audio. Specifically, the generative AI model is used to process the video and audio based on the feature data and emotion data.

[1137] Step 9:

[1138] The receiver's terminal presents the generated real-time video and audio to the user. The input is the generated video and audio, and the output is the video and audio presented to the user. Specifically, the system performs a process to play back the generated content using a display and speakers.

[1139] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1140] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1141] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1142] [Fourth embodiment]

[1143] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1144] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1145] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1146] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1147] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1148] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1149] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1150] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1151] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1152] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1153] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1154] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1155] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1156] The system of the present invention utilizes generative AI to generate video and audio in real time based on feature data in order to improve the efficiency of data transmission and reception in online conferences and video streaming.

[1157] Program processing overview

[1158] User connection and pre-transmission of data

[1159] The server receives profile images and voice samples of users when they connect to an online conference. During this process, users submit their own images and voice data introducing themselves, which are then stored on the server.

[1160] Extracting and sending feature data

[1161] The transmitting device captures the user's camera video and microphone audio in real time and extracts feature data from them, including facial landmarks (such as the positions of the eyes, nose, and mouth) and frequency spectrum characteristics of the voice.

[1162] The terminal compresses the extracted feature data and transmits it to the server as an efficient transmission packet.

[1163] Video Creation and Distribution

[1164] The server analyzes the feature data received from the sending device and transmits it to the receiving device in an appropriate format, allowing the generating AI to recreate the video and audio in real time.

[1165] The receiving device runs the generative AI using the feature data sent from the server to generate video and audio in real time, which is then displayed to the receiving user, providing an experience similar to that of a regular online meeting.

[1166] Specific examples

[1167] 1. Connection and Pre-Data Transmission

[1168] User A connects to the online conference, and the server receives the profile picture and self-introduction voice of User A. User A's terminal transmits this data to the server.

[1169] 2. Extracting and sending feature data

[1170] The transmitting terminal (User A's device) captures User A's video and audio in real time during the conference and extracts feature data such as facial position information and audio frequency characteristics. This feature data is efficiently compressed and sent to the server as a transmission packet.

[1171] 3. Video Creation and Distribution

[1172] The server analyzes the received feature data and sends it in an appropriate format to the receiving device (User B's device). The receiving device inputs the transmitted feature data into the generation AI, which generates real-time video and audio of User A. User B then watches the generated video and audio, enjoying a natural online meeting experience.

[1173] In this way, the system of the present invention can provide real-time video and audio while significantly reducing the amount of data transmitted and received and easing the strain on network resources, resulting in cost reduction and a reduction in the environmental impact.

[1174] The processing flow will be explained below.

[1175] Step 1:

[1176] A user connects to an online meeting.

[1177] The user launches the online meeting application and connects to the server.

[1178] Step 2:

[1179] The server receives the advance data.

[1180] The server receives a profile image and voice sample from the user, which includes a photo of the user's face and an audio introduction.

[1181] Step 3:

[1182] The user submits the advance data.

[1183] Users send their profile picture and voice sample from their device to the server, and this data is sent when the meeting is initially connected.

[1184] Step 4:

[1185] The device captures video and audio.

[1186] The sending device captures the user's camera video and microphone audio in real time, allowing conversations and actions to be recorded.

[1187] Step 5:

[1188] The device extracts the feature data.

[1189] The device extracts facial landmarks (such as the positions of the eyes, nose, and mouth) and gestures from the captured video, and also extracts frequency spectrum characteristics from the audio data.

[1190] Step 6:

[1191] The terminal compresses the feature data.

[1192] The terminal efficiently compresses the extracted feature data to generate transmission packets, which significantly reduces the amount of data.

[1193] Step 7:

[1194] The terminal transmits the feature data to the server.

[1195] The terminal transmits a transmission packet to the server in real time, and the transmission packet includes compressed feature data.

[1196] Step 8:

[1197] The server analyzes the feature data.

[1198] The server parses the received feature data and decodes it into the appropriate format, making the data usable by the receiving device.

[1199] Step 9:

[1200] The server transmits the data to the receiving terminal.

[1201] The server transmits the analyzed feature data to the receiving terminal, which then generates the data in real time.

[1202] Step 10:

[1203] The receiving device launches the generation AI.

[1204] The receiving device activates the generation AI based on the transmitted feature data and inputs the data, which starts the generation process.

[1205] Step 11:

[1206] Generative AI generates real-time video and audio.

[1207] The generative AI uses the feature data to generate real-time video and audio that replicates the movements and mouth movements of the sending user.

[1208] Step 12:

[1209] The receiving terminal displays the generated video and audio to the user.

[1210] The receiving terminal displays the generated video and audio to the user in real time, and the user can see and hear the video and audio of other participants just like in a regular online conference.

[1211] This system significantly reduces the amount of data sent and received, easing the load on network resources and resulting in cost savings and a reduced environmental impact.

[1212] Example 1

[1213] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1214] In online conferences and video streaming, sending and receiving real-time video and audio data requires a large amount of network resources, resulting in communication delays and quality degradation. Furthermore, existing technologies lack efficient methods for generating high-quality video and audio in real time. This results in poor user experience and increases costs and environmental impact.

[1215] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1216] In this invention, the server includes a means for a user to connect to an online conference and transmit a profile image and a voice sample, a transmitting terminal including means for extracting feature data including facial landmarks and voice frequency spectrum characteristics from real-time video and audio data, a means for compressing the feature data into a transmission packet, a means for transmitting the transmission packet to the server and for the server to convert the feature data into an appropriate format, a means for the server to transmit the converted data to a receiving terminal, and a means for the receiving terminal to generate real-time video and audio using a generative AI model based on the feature data. This significantly reduces the amount of data transmitted and received, enabling real-time generation of high-quality video and audio while minimizing communication delays.

[1217] "Users" are typical participants who use online meetings and video streaming services.

[1218] An "online conference" refers to a conference or meeting that takes place over the Internet in real time, sharing audio and video.

[1219] A "profile image" is a still image that a user uses to indicate their personal information or identity.

[1220] A "voice sample" is a short piece of voice data that a user provides to the server, used for self-introduction and identification.

[1221] The "transmitting terminal" is a device used by a user (for example, a smartphone or a personal computer) that captures video and audio using a camera and microphone.

[1222] "Feature data" are specific digital indices or characteristics (e.g., facial landmarks or frequency spectrum characteristics of voice) extracted from captured video and audio data.

[1223] "Landmarks" are characteristic points that indicate the position information of each part of the face (eyes, nose, mouth, etc.).

[1224] The "frequency spectrum characteristics" are data that indicate the strength of each frequency component in an audio signal.

[1225] A "server" is a central system that manages and processes data sent and received from sending and receiving terminals.

[1226] A "transmission packet" is a small unit of compressed data containing feature data that is transmitted over a network.

[1227] The "receiving terminal" is a device that receives the feature data sent from the server and reproduces real-time video and audio using a generation AI.

[1228] A "generative AI model" is an artificial intelligence technology that uses machine learning algorithms to generate video and audio from feature data.

[1229] "Compression" is a technique for reducing data size and is used to improve transmission efficiency.

[1230] "Real-time" refers to instantaneous processing and communication with little or no delay.

[1231] MODE FOR CARRYING OUT THE INVENTION

[1232] The system of the present invention utilizes a generative AI model to generate video and audio in real time based on feature data in order to improve the efficiency of data transmission and reception in online meetings and video streaming. A specific embodiment of this system is described below.

[1233] Hardware and software used

[1234] Hardware: A computer device with a compatible camera, microphone, and internet connection

[1235] Software: Generative AI models (e.g., GPT-3, DALL-E), libraries for video and audio capture (e.g., OpenCV, PyAudio)

[1236] A natural language description of the program

[1237] User connection and pre-transmission of data

[1238] A user opens an online conference app on their computer and connects to the server. The user then sends a profile picture and a self-introduction voice sample. During this process, the user follows the application's instructions to select a profile picture they have taken in advance, record a self-introduction voice using a microphone, and send it to the server. The server then stores the received data.

[1239] Real-time extraction and transmission of feature data

[1240] The sending device (e.g., User A's device) uses a camera and microphone to capture User A's video and audio in real time. Specifically, it uses a computer vision library such as OpenCV to detect facial landmarks and an audio analysis library such as PyAudio to analyze the frequency spectrum characteristics of the audio. These feature data are efficiently compressed and sent to the server as transmission packets.

[1241] Feature data analysis and resubmission

[1242] The server analyzes the feature data received from the sending device and converts the data into the appropriate format. This involves using generative AI models to calculate additional data needed to generate the video and audio. For example, it generates data to complete facial details and audio. The server then transmits the converted data to the receiving device.

[1243] Real-time video and audio generation

[1244] The receiving device (e.g., User B's device) receives the feature data and complementary data sent from the server, and the internal generative AI model is activated. For example, DALL-E generates a video reproducing User A's face based on the facial landmark data and complementary data, and simultaneously generates real-time audio based on the audio sample. The generated video and audio are output to the screen and speakers, just like a regular video conferencing app.

[1245] Specific examples

[1246] 1. Connection and Pre-Data Transmission

[1247] User A connects to the online conference, and the server receives the profile picture and self-introduction voice of User A. User A's terminal transmits this data to the server.

[1248] 2. Real-time extraction and transmission of feature data

[1249] The transmitting terminal (User A's device) captures User A's video and audio in real time during the conference and extracts feature data such as facial position information and audio frequency characteristics. This feature data is efficiently compressed and sent to the server as a transmission packet.

[1250] 3. Analysis and retransmission of feature data

[1251] The server analyzes the received feature data and sends it in an appropriate format to the receiving device of User B. User B's device inputs the transmitted feature data into the generation AI, which generates real-time video and audio of User A.

[1252] Prompt Sentence Examples

[1253] Data sent from User A's device to the server: "Send profile picture and voice sample"

[1254] The process in which the server sends feature data to User B's device: "Generate video and audio based on facial landmarks and audio spectrum characteristics."

[1255] This system significantly reduces the amount of data transmitted and received, alleviating pressure on network resources while providing high-quality real-time video and audio, resulting in cost savings and a lighter environmental impact.

[1256] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1257] Step 1:

[1258] User connection and pre-transmission of data

[1259] A user opens an online conference app on their computer and connects to the server. The user provides a profile image and a self-introduction voice sample as input, which is then sent to the server. The input image and voice data are stored on the server. For example, the user may upload a face photo taken with a camera or a voice recording with a microphone. This process is performed to collect initial data that will serve as the basis for subsequent feature data generation.

[1260] Step 2:

[1261] Real-time extraction and transmission of feature data

[1262] The sending terminal (user's device) uses a camera and microphone to capture the user's video and audio in real time. This becomes the input data. Specifically, a computer vision library such as OpenCV is used to detect facial landmarks (the positions of the eyes, nose, mouth, etc.), and an audio analysis library such as PyAudio is used to analyze the frequency spectrum characteristics of the audio. This feature data is compressed and sent to the server as a transmission packet. For example, this includes operations to obtain the position information of each part of the face from the camera video and calculate the strength of specific frequency components from the audio signal. The output is compressed feature data.

[1263] Step 3:

[1264] Feature data analysis and resubmission

[1265] The server receives feature data from the sending device as input. It analyzes this data and uses a generative AI model (e.g., GPT-3 or DALL-E) to calculate the necessary completion data for video and audio generation. Specifically, it generates completion data for facial details and audio. This data is then compressed again and sent to the receiving device in an appropriate format. For example, this may involve the generative AI completing fine details of the user's face to generate high-quality video data. The output is a packet containing the completion data and the original feature data.

[1266] Step 4:

[1267] Real-time video and audio generation

[1268] The receiving device (User B's device) receives the feature data and complementary data from the server as input. An internal generative AI model is activated and generates video and audio in real time based on this data. Specifically, DALL-E generates video recreating User A's face based on facial landmark data, and generates real-time audio based on audio samples. The output is the generated video and audio, which are displayed on User B's screen and played through the speakers.

[1269] In this way, the system of the present invention can provide high-quality video and audio in real time while significantly reducing the amount of data sent and received and easing the strain on network resources.

[1270] (Application example 1)

[1271] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1272] Conventional online conference systems and video streaming systems require large amounts of data for transmission and reception, resulting in problems such as strain on network resources and communication delays. Furthermore, real-time natural dialogue between customers and store staff in virtual stores requires the transmission and reception of high-quality video and audio. However, this data transmission depends on bandwidth and communication quality, which can potentially impair the customer experience. A system that can solve this issue and efficiently provide real-time video and audio is needed.

[1273] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1274] In this invention, the server includes a means for receiving and storing customer profile images and voice samples, a means for capturing video and audio of the virtual store clerk in real time, extracting and compressing feature data, and transmitting the data to the server, and a means for the receiving terminal to play back video and audio of the virtual store clerk in real time using artificial intelligence for generation, thereby enabling efficient use of network resources and providing high-quality real-time video and audio.

[1275] The "transmitting terminal" is a terminal that captures video and audio data and extracts feature data in real time.

[1276] "Feature data" is data extracted from video and audio data, such as facial landmarks and audio frequency spectrum characteristics.

[1277] The "means for compressing and forming transmission packets" refers to a means for efficiently compressing the extracted feature data and converting it into a packet format for transmission.

[1278] The "recipient's terminal" is a terminal that receives the feature data sent from the sender and generates video and audio in real time using artificial intelligence.

[1279] "Generative AI" is an AI technology that generates video and audio in real time based on feature data.

[1280] A "virtual store" is a store that exists in a virtual space and that users can access online to purchase goods and services.

[1281] A "profile picture" is a static image that a user provides when connecting to an online service.

[1282] "Audio sample" is audio data provided by a user when connecting to an online service.

[1283] A "virtual store clerk" is a virtual character that serves customers in a virtual store.

[1284] "Means for playing in real time" refers to means for instantly generating video and audio using artificial intelligence based on received feature data, and displaying and playing them to the user.

[1285] The system for implementing this invention includes a real-time customer service system in a virtual store. The system is composed of the following main components:

[1286] Server-side processing

[1287] When a customer visits the virtual store, the server receives and stores the customer's profile image and voice sample (e.g., self-introduction). The profile data is used for customer identification and data matching. Next, the server receives real-time video and audio data sent from the store clerk's terminal and extracts feature data from that data. Specifically, this includes facial landmarks and frequency spectrum characteristics of the voice. The extracted feature data is efficiently compressed and sent to the recipient's terminal.

[1288] Terminal side processing

[1289] The recipient's device (customer's device) receives the feature data sent from the server. Based on the received feature data, a generative AI model is used to generate video and audio of a virtual store clerk in real time. The generated video and audio are then played back to the customer in real time, allowing the customer to enjoy a natural communication experience, as if they were in a real store.

[1290] Hardware and software used

[1291] Server: A server with powerful CPU and GPU. Uses generative artificial intelligence models (e.g., OpenAI's GPT series).

[1292] Client device: Smartphone or head-mounted display (HMD). The OpenCV library is used to receive feature data and play back the generated video and audio.

[1293] Specific examples

[1294] Example prompt:

[1295] "When customers visit a virtual store, they communicate with a virtual associate in real time based on their visit."

[1296] This invention solves the data transmission and reception issues of conventional online conference systems and video streaming systems. It also enables real-time transmission of high-quality video and audio in virtual stores, improving the customer experience and enabling efficient use of network resources.

[1297] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1298] Step 1:

[1299] When a user visits the virtual store, the server receives and stores the user's profile picture and voice sample (e.g., self-introduction), which are used to identify the user and provide customized services.

[1300] Input: User's profile picture and voice sample.

[1301] Output: User profile data stored on the server.

[1302] How it works: The server stores user-submitted profile images and voice samples in a database.

[1303] Step 2:

[1304] The virtual store assistant uses a dedicated device to capture his or her own video and audio in real time, and the transmitting device extracts feature data such as facial landmarks and frequency spectrum characteristics from this data, then efficiently compresses and converts it into transmission packets.

[1305] Input: Real-time video and audio data.

[1306] Output: Outgoing packets containing compressed feature data.

[1307] How it works: The sending device analyzes the video and audio data to extract facial landmarks and audio frequency spectrum characteristics, then compresses and converts them into packets for transmission.

[1308] Step 3:

[1309] The server receives the transmitted packets from the sending terminal, analyzes them, and routes them appropriately to each receiving terminal.

[1310] Input: Transmission packets containing compressed feature data.

[1311] Output: Parsed feature data.

[1312] How it works: The server unpacks the received transmission packet and analyzes the feature data against the profile data stored in its internal database.

[1313] Step 4:

[1314] The receiver's device receives the feature data sent from the server and uses the generative AI model to generate real-time video and audio, which is then instantly played back to the user (customer).

[1315] Input: Feature data sent from the server.

[1316] Output: Real-time generated video and audio.

[1317] How it works: The recipient's device inputs the received feature data into a generative AI model, which then generates and plays back video and audio of a virtual store clerk based on this data.

[1318] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1319] The system of the present invention utilizes generative AI and an emotion engine to generate video and audio in real time and recognize emotions based on feature data in order to improve the efficiency of data transmission and reception in online meetings and video streaming.

[1320] Program processing overview

[1321] User connection and pre-transmission of data

[1322] The server receives profile images and voice samples of users when they connect to an online conference. During this process, users submit their own images and voice data introducing themselves, which are then stored on the server.

[1323] Extracting and sending feature data

[1324] The transmitting device captures the user's camera video and microphone audio in real time and extracts feature data from them, specifically facial landmarks (such as the positions of the eyes, nose, and mouth), gestures, and frequency spectrum characteristics from audio data.

[1325] The device recognizes the user's emotions based on the captured video and audio data using an emotion engine that analyzes emotions from facial expressions and voice tone.

[1326] The terminal compresses the extracted feature data and the recognized emotion data and transmits them to the server as efficient transmission packets.

[1327] Video Creation and Distribution

[1328] The server analyzes the feature and emotion data received from the sending device and transmits it to the receiving device in an appropriate format, allowing the generating AI to reproduce the video and audio in real time.

[1329] The receiving device uses the feature data and emotion data sent from the server to run the generative AI and generate video and audio in real time. The generated video and audio also reflect the recognized emotions, realizing natural communication.

[1330] Specific examples

[1331] 1. Connection and Pre-Data Transmission

[1332] User A connects to the online conference, and the server receives the profile picture and self-introduction voice of User A. User A's terminal transmits this data to the server.

[1333] 2. Extracting and sending feature data

[1334] The transmitting terminal (User A's device) captures the video and audio of User A in real time during the conference. From this data, feature data such as facial position information and audio frequency characteristics are extracted.

[1335] Furthermore, the emotion engine installed in User A's device recognizes User A's emotions from facial expressions and voice tone, analyzing emotions such as smiling or anger, for example.

[1336] The extracted feature data and emotion data are compressed and transmitted to the server as transmission packets.

[1337] 3. Video Creation and Distribution

[1338] The server analyzes the received feature data and emotion data and sends them in an appropriate format to the receiving device (User B's device). The receiving device inputs the transmitted data into the generation AI, which generates real-time video and audio of User A.

[1339] The generated video and audio also reflect User A's emotions, allowing User B to have a natural online meeting experience based on User A's facial expressions and tone.

[1340] In this way, the system of the present invention can provide natural communication through emotion recognition while significantly reducing the amount of data sent and received and easing the strain on network resources, resulting in cost reduction and a reduction in the environmental impact.

[1341] The processing flow will be explained below.

[1342] Step 1:

[1343] A user connects to an online meeting.

[1344] The user launches an online conference application and connects to the server by entering a URL or ID to connect to a specific conference room.

[1345] Step 2:

[1346] The server receives the advance data.

[1347] Upon initial connection, the server receives a profile image and voice sample from the user, including a photo of the user's face and a voice clip introducing themselves.

[1348] Step 3:

[1349] The user submits the advance data.

[1350] The user device sends a facial photo taken with a camera and a voice sample recorded with a microphone to the server, and this pre-data is used for user identification and initial setup.

[1351] Step 4:

[1352] The device captures video and audio.

[1353] The sending device captures the user's camera video and microphone audio in real time, recording the video and audio data frame by frame.

[1354] Step 5:

[1355] The device extracts the feature data.

[1356] The device detects facial landmarks (such as the eyes, nose, and mouth) and gestures from the captured video. Frequency spectrum characteristics are also extracted from the audio data. This allows the device to accurately understand the user's movements and speech while reducing the amount of data.

[1357] Step 6:

[1358] The device runs the emotion engine.

[1359] The device's built-in emotion engine recognizes the user's emotions from captured video and audio data, using a dedicated algorithm to identify emotions such as joy, anger, sadness, and happiness from facial expressions and voice tone.

[1360] Step 7:

[1361] The device compresses the feature data and emotion data.

[1362] The device efficiently compresses the extracted feature data and emotion data and sends them as packets, which are small in size and minimize network load.

[1363] Step 8:

[1364] The terminal transmits a transmission packet to the server.

[1365] The device sends packets containing compressed feature data and emotion data to the server in real time, a process that is fast and causes almost no delay.

[1366] Step 9:

[1367] The server analyzes the feature data and emotion data.

[1368] The server analyzes the received transmission packets and decodes the feature data and emotion data, converting them into a format that can be used by the receiving terminal.

[1369] Step 10:

[1370] The server transmits the data to the receiving terminal.

[1371] The server then transmits the analyzed feature data and emotion data to the receiving device, which receives this data and uses it as input for the generation AI.

[1372] Step 11:

[1373] The receiving device launches the generation AI.

[1374] The receiving device inputs the transmitted feature data and emotion data into the generation AI, initiating the real-time video and audio generation process.

[1375] Step 12:

[1376] Generative AI generates real-time video and audio.

[1377] The generative AI generates real-time video and audio of the user from feature data, and also reflects the user's facial expressions and vocal intonation based on emotional data.

[1378] Step 13:

[1379] The receiving terminal displays the generated video and audio to the user.

[1380] The receiving terminal displays the generated video and audio to the user in real time, allowing the user to view natural video and audio that reflects the emotions of the sending user.

[1381] This system significantly reduces the amount of data sent and received, alleviating the burden on network resources while providing natural communication through emotion recognition, resulting in cost savings and a lighter environmental impact.

[1382] Example 2

[1383] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1384] In online meetings and video streaming, it is important to generate high-quality video and audio in real time to improve the efficiency of data transmission and reception. It is also necessary to recognize user emotions and realize natural communication.

[1385] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a sending terminal including means for extracting feature data from real-time video and audio data, means for recognizing a user's emotion based on the feature data, means for compressing the feature data and emotion data into transmission packets, means for transmitting the transmission packets to a receiving terminal, and means for the receiving terminal to generate real-time video and audio using a generation AI based on the feature data and emotion data. This enables efficient data transmission and reception and natural communication.

[1386] A "transmitting terminal" is a device that captures video and audio data in real time, extracts feature data from the data, and transmits the data to a server.

[1387] "Feature data" refers to information such as facial landmarks and frequency spectrum characteristics extracted from the user's video and audio data.

[1388] "Emotion data" is information about emotions analyzed from the user's video and audio by an emotion recognition engine.

[1389] A "transmission packet" is a data packet that contains compressed feature data and emotion data.

[1390] The "recipient's terminal" is a device that receives the transmission packets sent from the server and generates video and audio in real time using generation AI.

[1391] "Generative AI" is an artificial intelligence model that generates video and audio in real time based on feature data and emotional data.

[1392] An "emotion recognition engine" is software that analyzes emotions from a user's video and audio data and outputs them as emotional data.

[1393] A "machine learning model" is a model that uses algorithms to learn patterns and regularities from data and make inferences and predictions.

[1394] The system of the present invention improves the efficiency of data transmission and reception in online meetings and video streaming, and utilizes generative AI and an emotion engine to generate video and audio in real time and recognize emotions based on feature data.

[1395] User connection and pre-transmission of data

[1396] When a user connects to an online conference, the user sends their profile picture and self-introduction audio data from their device to the server, which then receives and stores this data for each user.

[1397] Extracting and sending feature data

[1398] The sending device captures the user's camera video and microphone audio in real time during an online conference. From this captured data, the device extracts frequency spectrum characteristics from facial landmarks (such as the positions of the eyes, nose, and mouth), gestures, and voice data. Furthermore, it uses an emotion engine to recognize the user's emotions and generate emotion data.

[1399] The extracted feature data and emotion data are compressed and sent to the server as transmission packets. This compression and transmission process is important for enabling efficient communication.

[1400] Video Creation and Distribution

[1401] The server analyzes the transmitted feature data and emotion data and transmits it to the recipient's device in the appropriate format. This data is formatted so that the generative AI can reproduce the video and audio in real time. The recipient's device inputs the data transmitted from the server into the generative AI model, generating the video and audio in real time. The generated video and audio reflect the user's emotions, enabling natural communication.

[1402] Specific examples

[1403] 1. Connection and Pre-Data Transmission

[1404] When user A connects to an online conference, the terminal sends a profile image (profile.jpg) and self-introduction audio data (intro.mp3) to the server.

[1405] The server receives this data, associates it with User A's ID, and stores it in a database.

[1406] 2. Extracting and sending feature data

[1407] The transmitting device captures the video and audio of User A in real time. From this captured data, facial landmarks (eyes (x1, y1), nose (x2, y2), mouth (x3, y3)) and audio frequency characteristics (frequency band Hz, amplitude dB) are extracted.

[1408] The emotion engine recognizes emotions from user A's facial expressions and tone of voice, and generates emotion data for "joy," for example.

[1409] The extracted feature data and emotion data are compressed and transmitted to the server as transmission packets.

[1410] 3. Video Creation and Distribution

[1411] The server analyzes the received feature data and emotion data and transmits them to the receiving terminal (user B's terminal) in an appropriate format.

[1412] The receiving device inputs the transmitted feature data and emotion data into the generative AI model, generating video and audio of User A in real time. Because the generated video and audio also reflect User A's emotions, User B can enjoy a natural online meeting experience based on User A's facial expressions and tone.

[1413] Prompt Sentence Examples

[1414] "Analyze the camera video and microphone audio of User A, and generate video and audio in real time based on the following feature data. The feature data is as follows:

[1415] Facial landmarks: eyes(x,y), nose(x,y), mouth(x,y)

[1416] Audio frequency spectrum: frequency band (Hz), amplitude (dB)

[1417] Emotion data: smile, anger

[1418] The generated video and audio should reflect the perceived emotions of User A.

[1419] In this way, the system of the present invention can provide efficient data transmission and reception and natural communication, thereby reducing costs and environmental impact.

[1420] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1421] Step 1:

[1422] When a user connects to an online conference, the user sends a profile image and self-introduction voice data from the terminal to the server. The server receives the user's profile image and self-introduction voice data and stores them in a database for each user.

[1423] Input: User's profile image (e.g., profile.jpg) and self-introduction audio data (e.g., intro.mp3)

[1424] Output: User profile image and self-introduction audio data stored in the server database

[1425] Specific operation: The user selects a profile picture and voice data on the device and clicks the "Send" button. The server associates the received data with the user ID and stores it in the database.

[1426] Step 2:

[1427] The sending device captures the user's camera video and microphone audio in real time during the online conference, so that the user's current video and audio are input to the device.

[1428] Input: Real-time captured video and audio data

[1429] Output: Video frames and audio samples captured in real time

[1430] Specific operation: The camera captures the user's image and the microphone records the user's voice. These data are stored in the device's memory.

[1431] Step 3:

[1432] The transmitting terminal extracts facial landmarks (such as the positions of the eyes, nose, and mouth) and gestures from the captured video data, and frequency spectrum characteristics from the audio data.

[1433] Input: Video frames and audio samples captured in real time

[1434] Output: Extracted facial landmarks (e.g., eye position (x1, y1), nose position (x2, y2), mouth position (x3, y3)), frequency spectrum characteristics (e.g., frequency band Hz, amplitude dB)

[1435] How it works: The image processing algorithm analyzes the video data and detects facial features, while the audio analysis algorithm analyzes the frequency characteristics of the audio data.

[1436] Step 4:

[1437] The transmitting terminal uses an emotion engine to recognize emotions from the user's facial expressions and tone of voice, and generates emotion data.

[1438] Input: Video frames and audio samples captured in real time

[1439] Output: Recognized emotion data (e.g., "joy", "anger")

[1440] Specific operation: The emotion recognition engine analyzes facial expressions and voice to identify the user's emotional state, resulting in emotional data.

[1441] Step 5:

[1442] The transmitting terminal compresses the extracted feature data and emotion data and transmits them as transmission packets to the server.

[1443] Input: extracted feature data and emotion data

[1444] Output: Compressed outgoing packets

[1445] Specific operation: The data compression algorithm compresses the feature data and emotion data to generate a transmission packet, which the device then sends to the server.

[1446] Step 6:

[1447] The server analyzes the transmitted feature data and emotion data and transmits them to the receiving terminal in an appropriate format.

[1448] Input: Outgoing packets

[1449] Output: Data reformatted to the appropriate format

[1450] Specific operation: The server decompresses and analyzes the transmitted packets, reformats the analysis results into a format usable by the receiving terminal, and transmits them.

[1451] Step 7:

[1452] The receiving terminal inputs the feature data and emotion data sent from the server into the generative AI model and generates video and audio in real time.

[1453] Input: Feature data and emotion data sent from the server

[1454] Output: Generated real-time video and audio

[1455] How it works: A prompt is input into the generative AI model, which generates video and audio in real time, with the recognized emotion reflected in the generated video and audio.

[1456] The above processing steps enable efficient data transmission and reception and natural communication in online conferences and video streaming.

[1457] (Application example 2)

[1458] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1459] Online conferences and video streaming services that utilize real-time video and audio data require the transmission and reception of large amounts of data, resulting in increased communication costs and network latency. Furthermore, it is difficult to accurately convey users' emotions, which can impede natural communication. This creates a demand for more efficient and natural communication methods.

[1460] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1461] In this invention, the server includes a sending terminal that extracts feature data from real-time video and audio data, a means for compressing the feature data and emotion data into transmission packets, a means for transmitting the transmission packets to a receiving terminal, and a means for the receiving terminal to generate real-time video and audio using a generation AI based on the feature data and emotion data. This makes it possible to generate natural video and audio that reflects the user's emotions while significantly reducing the amount of data sent and received.

[1462] The "transmitting terminal" is a device where a user generates video and audio data in real time and extracts this as feature data and emotion data.

[1463] "Feature data" is analytical information such as facial landmarks and audio frequency characteristics extracted from real-time video and audio data.

[1464] "Emotion data" is information obtained as a result of analyzing the user's emotional state based on the extracted feature data.

[1465] A "transmission packet" is a data packet containing compressed feature data and emotion data, which is transmitted over a network.

[1466] The "recipient terminal" is a device that receives the transmitted packets and runs the generation AI to generate video and audio in real time.

[1467] "Generative AI" is an artificial intelligence technology that generates video and audio in real time based on received feature data and emotional data.

[1468] A "specialized algorithm" is a program that uses specific processing techniques to extract feature data and emotion data from video and audio data.

[1469] A "machine learning model" is a trained model that, as part of generative AI, generates video and audio in real time based on incoming data.

[1470] MODE FOR CARRYING OUT THE INVENTION

[1471] An object of the present invention is to provide a system that improves the efficiency of data transmission and reception in online conferences and live streaming services, and realizes natural communication that reflects the user's emotions. The following describes in detail an embodiment of the present invention.

[1472] Program processing overview

[1473] The system of the present invention consists of three main components: a sending terminal, a server, and a receiving terminal.

[1474] Sending device

[1475] The transmitting device captures real-time video and audio and extracts feature and emotion data from it. Specifically, it acquires video and audio using hardware such as a camera and microphone, and analyzes facial landmarks and audio frequency characteristics using a dedicated algorithm. Based on the results of this analysis, the emotion engine recognizes the user's emotions. The extracted feature and emotion data are compressed and sent to the server as a transmission packet.

[1476] server

[1477] The server manages and analyzes the feature data and emotion data received from the sending device and sends it to the receiving device. To ensure efficient data transmission and reception, the server converts the feature data and emotion data into an appropriate format and prepares it in a form that is easy for the generation AI to process before sending it.

[1478] Recipient's device

[1479] The receiver's device runs a generation AI based on the data received from the server to generate real-time video and audio. The generation AI uses a machine learning model to reproduce the feature data and emotion data compressed for transmission in real time, generating natural-looking video and audio that also reflects the user's emotions.

[1480] Hardware and Software Used

[1481] Sending device: camera, microphone, specialized algorithms (e.g., facial landmark detection algorithms)

[1482] Emotion engine: Software for recognizing emotions from audio and video (e.g., EmotionEngine)

[1483] Server: Server software that manages and analyzes data transmission and reception

[1484] Recipient device: Generative AI model (e.g., a generative AI using a specific machine learning model)

[1485] Data processing and calculation

[1486] Processing on the sending device: Facial landmarks and audio frequency characteristics are extracted from video and audio captured in real time, and emotions are recognized using an emotion engine.

[1487] Data compression and transmission: The extracted feature data and emotion data are compressed and sent to the server as a transmission packet.

[1488] Processing on the server: The received data is analyzed, converted into a format that is easy for the generating AI to process, and sent to the recipient's device.

[1489] Processing on the recipient's device: Using generative AI, video and audio are generated in real time based on the transmitted feature data and emotion data.

[1490] Specific examples

[1491] When a live streamer broadcasts live on their smartphone, they register a profile picture and audio sample in advance. During the live broadcast, the sending device captures video and audio in real time, extracts feature data and emotional data, and sends it to a server. The receiving device runs a generative AI based on the received data to generate video and audio in real time. This method makes it possible to provide viewers with high-quality video and audio with low latency.

[1492] Example prompt sentence:

[1493] "Generate video and audio in real time based on the user's profile picture and voice sample. Recognize the user's emotions from facial expressions and tone of voice."

[1494] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1495] Step 1:

[1496] The sending device captures video and audio in real time from the user's camera and microphone. The input is the camera video and microphone audio, and the output is the captured video frames and audio samples of these data. Specifically, it performs a process to acquire the data stream from the camera and microphone.

[1497] Step 2:

[1498] The transmitting device detects facial landmarks (such as the eyes, nose, and mouth) from the captured video and extracts frequency spectrum characteristics from the audio. The input is video frames and audio samples, and the output is facial landmark data and audio characteristic data. Specifically, it runs a facial landmark detection algorithm and an audio spectrum analysis algorithm.

[1499] Step 3:

[1500] The transmitting device uses an emotion engine to recognize the user's emotion based on the detected facial landmarks and voice characteristic data. The input is facial landmark data and voice characteristic data, and the output is recognized emotion data. Specifically, the emotion engine is executed to analyze the user's emotion from their facial expression and voice tone.

[1501] Step 4:

[1502] The transmitting terminal compresses the extracted feature data and emotion data into transmission packets. The input is facial landmark data, voice characteristic data, and emotion data, and the output is compressed transmission packets. Specifically, the data is efficiently packetized using a data compression algorithm.

[1503] Step 5:

[1504] The sending terminal sends compressed transmission packets to the server. The input is the transmission packet, and the output is the data sent to the server. Specifically, the data is sent to the server using a network protocol.

[1505] Step 6:

[1506] The server decompresses the received transmission packets and analyzes the feature data and emotion data. The input is the received transmission packets, and the output is the decompressed feature data and emotion data. Specifically, the server executes a process to decompress the packets using a data decompression algorithm.

[1507] Step 7:

[1508] The server transmits the feature data and emotion data to the recipient's terminal. The input is the decompressed feature data and emotion data, and the output is the data transmitted to the recipient's terminal. Specifically, the data is transmitted using a network protocol.

[1509] Step 8:

[1510] The receiver's device inputs the received feature data and emotion data into the generative AI to generate real-time video and audio. The input is the received feature data and emotion data, and the output is the generated video and audio. Specifically, the generative AI model is used to process the video and audio based on the feature data and emotion data.

[1511] Step 9:

[1512] The receiver's terminal presents the generated real-time video and audio to the user. The input is the generated video and audio, and the output is the video and audio presented to the user. Specifically, the system performs a process to play back the generated content using a display and speakers.

[1513] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1514] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1515] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1516] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1517] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1518] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1519] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1520] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1521] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1522] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1523] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1524] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1525] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1526] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1527] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1528] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1529] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1530] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1531] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1532] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1533] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1534] The following is further disclosed regarding the above embodiment.

[1535] (Claim 1)

[1536] A transmitting terminal includes means for extracting feature data from real-time video and audio data;

[1537] means for compressing the characteristic data into a transmission packet;

[1538] means for transmitting the transmission packet to a receiver terminal;

[1539] A receiver terminal uses a generation AI to generate real-time video and audio based on the feature data;

[1540] A system including:

[1541] (Claim 2)

[1542] 2. The system of claim 1, wherein the transmitting terminal uses a dedicated algorithm as a means for extracting the characteristic data.

[1543] (Claim 3)

[1544] The system of claim 1, wherein the receiver's terminal uses a machine learning model as a means for generating video and audio using the generation AI.

[1545] "Example 1"

[1546] (Claim 1)

[1547] means for users to connect to an online conference and submit a profile picture and voice sample;

[1548] A transmitting terminal includes a means for extracting feature data including facial landmarks and frequency spectrum characteristics of voice from real-time video and audio data;

[1549] means for compressing the characteristic data into a transmission packet;

[1550] means for transmitting the transmission packet to a server, and the server converting the feature data into an appropriate format;

[1551] A means for the server to transmit the converted data to a recipient's terminal;

[1552] A receiver terminal generates real-time video and audio using a generation AI model based on the feature data;

[1553] A system including:

[1554] (Claim 2)

[1555] 2. The system of claim 1, wherein the transmitting terminal uses a computer vision library and a voice analysis library as means for extracting feature data.

[1556] (Claim 3)

[1557] The system of claim 1, wherein the receiver's terminal uses a machine learning algorithm as a means for generating video and audio using a generative AI model.

[1558] "Application Example 1"

[1559] (Claim 1)

[1560] A transmitting terminal includes means for extracting feature data from real-time video and audio data;

[1561] means for compressing the characteristic data into a transmission packet;

[1562] means for transmitting the transmission packet to a receiver terminal;

[1563] A receiver terminal uses artificial intelligence to generate real-time video and audio based on the feature data;

[1564] means for receiving and storing profile images and voice samples of customers accessing the virtual store;

[1565] a means for capturing video and audio of the virtual store clerk in real time, extracting and compressing feature data, and transmitting the data to a server;

[1566] A means for the receiving terminal to play back the video and audio of the virtual store clerk in real time using a generating artificial intelligence;

[1567] A system including:

[1568] (Claim 2)

[1569] 2. The system of claim 1, wherein the transmitting terminal uses a dedicated algorithm as a means for extracting the characteristic data.

[1570] (Claim 3)

[1571] The system of claim 1, wherein the receiver's terminal uses a machine learning model as a means for generating video and audio using generative artificial intelligence.

[1572] "Example 2: Combining Emotion Engines"

[1573] (Claim 1)

[1574] A transmitting terminal includes means for extracting feature data from real-time video and audio data;

[1575] means for recognizing a user's emotion based on the feature data;

[1576] means for compressing the feature data and emotion data into transmission packets;

[1577] means for transmitting the transmission packet to a receiver terminal;

[1578] A receiver's terminal uses a generation AI to generate real-time video and audio based on the feature data and emotion data;

[1579] A system including:

[1580] (Claim 2)

[1581] 2. The system of claim 1, wherein the sending terminal uses a dedicated algorithm and emotion recognition engine as a means for extracting feature data.

[1582] (Claim 3)

[1583] The system of claim 1, wherein the receiver's terminal uses a machine learning model as a means for generating video and audio using the generation AI.

[1584] "Application example 2 when combining emotion engines"

[1585] (Claim 1)

[1586] A transmitting terminal includes means for extracting feature data from real-time video and audio data;

[1587] means for compressing the feature data and emotion data into transmission packets;

[1588] means for transmitting the transmission packet to a receiver terminal;

[1589] A receiver's terminal uses a generation AI to generate real-time video and audio based on the feature data and emotion data;

[1590] A system including:

[1591] (Claim 2)

[1592] 2. The system of claim 1, wherein the sending terminal uses a dedicated algorithm as a means for extracting feature data and emotion data.

[1593] (Claim 3)

[1594] The system of claim 1, wherein the receiver's terminal uses a machine learning model as a means for generating video and audio using the generation AI. [Explanation of symbols]

[1595] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A transmitting terminal includes means for extracting feature data from real-time video and audio data; means for compressing the characteristic data into a transmission packet; means for transmitting the transmission packet to a receiver terminal; A receiver terminal uses a generation AI to generate real-time video and audio based on the feature data; A system including:

2. 2. The system of claim 1, wherein the transmitting terminal uses a dedicated algorithm as a means for extracting the characteristic data.

3. The system of claim 1, wherein the receiver's terminal uses a machine learning model as a means for generating video and audio using the generation AI.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A