System

A generative AI-based system converts presentation audio and video data to enhance workplace communication by reducing status-related barriers, allowing for more candid discussions.

JP2026017304APending Publication Date: 2026-02-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024118086
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-02-04

AI Technical Summary

Technical Problem

Differences in status within the workplace hinder effective communication during presentations by senior executives, particularly in remote work settings, making it difficult for subordinates and colleagues to express honest opinions.

Method used

A system that utilizes generative AI to convert presentation audio and video data into alternative formats, synthesizing them to create a more inclusive environment for understanding and expression of opinions.

Benefits of technology

The system facilitates candid opinions and improves workplace communication by reducing barriers caused by status differences, enabling more relaxed and constructive discussions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017304000001_ABST
    Figure 2026017304000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving uploaded presentation audio and video data; audio conversion means for converting the presentation audio data into other audio data; video conversion means for converting the presentation video data into other video data; and means for combining and outputting the converted audio and video data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] During presentations by senior executives in the workplace, subordinates and colleagues often find it difficult to express their honest opinions due to differences in their positions. This situation impedes communication in the workplace and hinders constructive discussion and exchange of opinions. Furthermore, in today's world where remote work is on the rise and face-to-face dialogue is difficult, this impact is even more pronounced. To solve this problem, technology is needed to convert presentations by senior executives into presentations for colleagues and subordinates. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for receiving uploaded presentation audio and video data, an audio conversion means for converting the presentation audio data into other audio data, a video conversion means for converting the presentation video data into other video data, and a means for synthesizing and outputting the converted audio and video data. This system uses generative AI to convert the audio and video data, thereby creating an environment where listeners of the presentation can understand the content without feeling discriminated against. This makes it easier for subordinates and colleagues to express their honest opinions, improving communication within the workplace.

[0006] "Uploading" is the act of a user sending data over a computer or network to a remote system, such as a server.

[0007] A "presentation" is a means of explanation or presentation that combines visual and audio materials to convey information or ideas to an audience.

[0008] "Audio data" is a data format for recording, storing, and transmitting sound in digital form.

[0009] "Video data" is a data format for recording, storing, and transmitting visual information in digital form.

[0010] "Means for receiving" refers to a function or device that allows a server or computer system to receive data sent from a user.

[0011] "Audio transformation means" means a system, device or algorithm for transforming original audio data into different audio data.

[0012] "Video transformation means" means a system, device, or algorithm for transforming original video data into different video data.

[0013] "Means for synthesizing" means a system, device or algorithm that combines the converted audio and video data into a single integrated data set.

[0014] "Output means" refers to a system or device that stores, displays, plays, or transmits the converted or synthesized data to provide it to the user.

[0015] "Generative AI" is an artificial intelligence technology that uses deep learning and other advanced algorithms to create and transform new data.

[0016] "Differences in status" is a concept that refers to differences in authority and responsibility based on job title or hierarchy, and this can sometimes become a barrier to communication.

[0017] "Environment" is a broad concept that refers to the external circumstances or conditions that affect a user's activities. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The present invention relates to a system including means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, and means for synthesizing and outputting the converted audio and video data. The system of the present invention utilizes generative AI to encourage candid opinions and improve communication in the workplace.

[0040] Program processing

[0041] 1. User upload of data

[0042] Users upload the presentation's audio and video files to the system, which is done using a common web interface.

[0043] 2. Data preprocessing by the server

[0044] The server receives the uploaded file and separates the audio and video data. The audio and video are saved as separate files, and preprocessing such as noise reduction and resolution conversion is performed as needed.

[0045] 3. Data transformation by the server

[0046] The server uses deep learning models to convert audio and video data: the audio conversion model converts the original audio data into other audio data, and the video conversion model converts the original video data into other video data.

[0047] 4. Data synthesis and output by the server

[0048] The converted audio and video data are then combined by the server and output as a new presentation file, which is then provided to the user.

[0049] Specific examples

[0050] For example, imagine a project manager at a company receives an important presentation from an executive, saved in a file called "original_presentation.mp4."

[0051] 1. The project manager (user) uploads "original_presentation.mp4" to the system.

[0052] 2. The server receives this file and separates the audio and video into "original_audio.mp3" and "original_video.mp4" respectively.

[0053] 3. The server applies the voice conversion model to convert "original_audio.mp3" into the voice of the project manager's colleague, resulting in the audio file "converted_audio.mp3."

[0054] 4. The server then applies the video conversion model to convert "original_video.mp4" into a video showing the colleague's face, resulting in a video file called "converted_video.mp4."

[0055] 5. Finally, the server combines "converted_audio.mp3" and "converted_video.mp4" to create a new file called "final_presentation.mp4."

[0056] 6. The project manager (user) downloads this "final_presentation.mp4" and watches the presentation with the faces and voices of their colleagues.

[0057] This allows project managers, unlike executives, to listen to presentations in a more relaxed state and more easily express their opinions frankly. This is expected to lead to improved communication and constructive exchange of opinions within the workplace. In this way, the system of the present invention reduces communication barriers caused by differences in position and provides an effective means for improving the work environment.

[0058] The processing flow will be explained below.

[0059] Step 1:

[0060] Users upload presentation audio and video files to the system. This process is done using a web interface, where users select files and send them to the server, where they are received and stored.

[0061] Step 2:

[0062] The server receives the uploaded file and separates the audio and video data. The server then reads the received video file, extracts the audio, and saves it as a separate file. Specifically, the audio is saved as "original_audio.mp3" and the video is saved as "original_video.mp4."

[0063] Step 3:

[0064] The server pre-processes the audio data by denoising and normalising it, and also adjusts the resolution and frame rate of the video data as needed, preparing it for a more efficient and accurate conversion process.

[0065] Step 4:

[0066] The server uses a voice conversion model to convert the original audio data into a target voice. In this case, it uses a deep learning model to convert it into the voice of the project manager's colleague. The converted result is saved as "converted_audio.mp3."

[0067] Step 5:

[0068] The server uses a video conversion model to convert the original video data into a target video, in this case the face of the project manager's colleague, using a deep learning model. The converted video data is saved as "converted_video.mp4."

[0069] Step 6:

[0070] The server combines the converted audio and video data. Specifically, it merges "converted_audio.mp3" and "converted_video.mp4" to generate a new presentation file. This new file is saved as "final_presentation.mp4."

[0071] Step 7:

[0072] Users download "final_presentation.mp4" from the system. They can play this file and watch the presentation being given by their colleagues, with their faces and voices. This process creates an environment where participants feel more comfortable listening to a presentation from a senior executive and can express their opinions more frankly.

[0073] Through these steps, the system of the present invention can provide users with an effective means of reducing status differences and improving communication within the workplace.

[0074] Example 1

[0075] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0076] Current presentations rely on the presenter's specific audio and video, making it difficult for viewers to evaluate the content unbiased. Furthermore, when it comes to smooth communication in the workplace, differences in perspective can make it difficult for people to express their opinions frankly. Conventional systems are limited in the technology they use, making it difficult to efficiently convert and synthesize audio and video.

[0077] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0078] In this invention, the server includes means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, means for converting the audio data and video data using a deep learning model, means for performing noise reduction and resolution conversion as preprocessing, and means for synthesizing and outputting the converted audio and video data. This reduces communication barriers due to differences in positions in the workplace and enables constructive exchange of opinions.

[0079] "Upload" is the process by which a user submits data to the system for storage.

[0080] "Presentation audio data" is digital data of the audio used during a presentation.

[0081] "Presentation video data" refers to digital data of the video used during a presentation.

[0082] "Audio conversion means" refers to a method or device for converting original audio data into audio data of a different format or content.

[0083] "Video conversion means" refers to a method or device for converting original video data into video data of a different format or content.

[0084] A "deep learning model" is a machine learning model that uses a multi-layer neural network to convert and analyze data.

[0085] "Preprocessing" refers to the process of performing basic processing such as noise removal and resolution conversion before converting the data.

[0086] "Denoising" is the process of removing unwanted noise from audio data.

[0087] "Resolution conversion" is the process of changing the resolution of video data.

[0088] "Synthesis" is the process of combining multiple pieces of data (audio data or video data) into one piece of data.

[0089] "Output" refers to providing processed data in a form usable by a user.

[0090] The present invention is a system that receives audio and video data of a presentation, converts it, and outputs it. The system processes presentation data uploaded by a user and converts the audio and video data into other formats using a deep learning model, enabling more effective communication.

[0091] First, a user uploads the presentation audio and video files to the system using the web interface. For example, a user may upload a file called "original_presentation.mp4." To do this, the user accesses the specified URL through a web browser and clicks the upload button.

[0092] Next, the server receives the uploaded file and separates the audio and video data. Specifically, it uses a tool such as FFmpeg to save the audio and video as separate files. It then applies a noise reduction filter to the separated audio data and performs resolution conversion on the video data. For example, the audio data is saved as "original_audio.mp3" and the video data as "original_video.mp4."

[0093] The server then uses a deep learning model to convert the audio and video data. It uses an audio conversion model (for example, a model using TensorFlow or PyTorch) to convert the audio data and generate a new audio file called "converted_audio.mp3". Similarly, it uses a video conversion model to generate a new video file called "converted_video.mp4".

[0094] The converted audio and video data are then combined by the server and output as the final presentation file, "final_presentation.mp4." Using a media processing library such as FFmpeg, the audio and video data are combined into a single file. Users can then download this new file and view it in their own environment.

[0095] As a concrete example, consider the case where a project manager at a company receives an important presentation from an executive. This presentation is saved in "original_presentation.mp4." The project manager (user) uploads this file to the system. The server receives this file and separates the audio and video. After noise removal and resolution conversion, it applies audio conversion models and video conversion models to generate new data for each. Finally, it is provided to the user as "final_presentation.mp4."

[0096] Example prompt sentence:

[0097] "Please upload your presentation file. Filename: original_presentation.mp4"

[0098] "Separating audio and video data. Please wait a moment."

[0099] "Converting audio data. Progress: 40%"

[0100] "Converting video data. Progress: 70%"

[0101] "Synthesizing converted data. Progress: 90%"

[0102] "Conversion complete. Download your file. Filename: final_presentation.mp4"

[0103] In this way, the system of the present invention uses deep learning technology and various media processing tools to effectively convert and synthesize presentation data uploaded by users, thereby improving communication within the workplace.

[0104] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0105] Step 1: User uploads data

[0106] A user uploads a presentation's audio and video files to the system using a web interface. The input is a file called "original_presentation.mp4," which is sent to the server. Specifically, the user opens a web browser, accesses the specified URL, clicks the "Choose File" button, selects the desired file, and presses the "Upload" button.

[0107] input:

[0108] Upload file: original_presentation.mp4

[0109] output:

[0110] The file is sent to the server

[0111] Step 2: Server receives and separates data

[0112] The server receives the uploaded "original_presentation.mp4" and uses FFmpeg to separate the audio and video data. Specifically, it saves the received file in a specific directory and executes the FFmpeg command to generate the audio data "original_audio.mp3" and the video data "original_video.mp4."

[0113] input:

[0114] Upload file: original_presentation.mp4

[0115] output:

[0116] Separated audio data: original_audio.mp3

[0117] Separated video data: original_video.mp4

[0118] Step 3: Preprocessing the data on the server

[0119] The server performs preprocessing on the separated audio and video data, applying an audio filter to remove noise and converting the resolution. Specifically, it uses the FFmpeg command to remove noise from the audio data (e.g., ffmpeg -i original_audio.mp3 -af "highpass=f=200, lowpass=f=3000" cleaned_audio.mp3) and converts the resolution of the video data (e.g., ffmpeg -i original_video.mp4 -vf scale=1280:720 resized_video.mp4).

[0120] input:

[0121] Audio data: original_audio.mp3

[0122] Video data: original_video.mp4

[0123] output:

[0124] Noise-removed audio data: cleaned_audio.mp3

[0125] Resolution converted video data: resized_video.mp4

[0126] Step 4: Data transformation by the server

[0127] The server converts audio and video data using a deep learning model. Specifically, it inputs audio data into an audio conversion model using TensorFlow or PyTorch, generating a converted audio file called "converted_audio.mp3." Similarly, it inputs video data into a video conversion model and generates a converted video file called "converted_video.mp4" (e.g., python audio_conversion.py --input cleaned_audio.mp3 --output converted_audio.mp3 and python video_conversion.py --input resized_video.mp4 --output converted_video.mp4).

[0128] input:

[0129] Preprocessed audio data: cleaned_audio.mp3

[0130] Preprocessed video data: resized_video.mp4

[0131] output:

[0132] Converted audio data: converted_audio.mp3

[0133] Converted video data: converted_video.mp4

[0134] Step 5: Data synthesis by the server

[0135] The server combines the converted audio and video data to generate the final presentation file, "final_presentation.mp4." Specifically, it uses FFmpeg to execute a command to combine the audio and video into a single file (e.g., ffmpeg -i converted_video.mp4 -i converted_audio.mp3 -c:v copy -c:a aac final_presentation.mp4).

[0136] input:

[0137] Converted audio data: converted_audio.mp3

[0138] Converted video data: converted_video.mp4

[0139] output:

[0140] Final presentation file: final_presentation.mp4

[0141] Step 6: User downloads final file

[0142] The user downloads the generated final presentation file, “final_presentation.mp4,” via the web interface by clicking the download link provided by the system.

[0143] input:

[0144] Final presentation file: final_presentation.mp4

[0145] output:

[0146] Downloaded file: final_presentation.mp4

[0147] (Application example 1)

[0148] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0149] With the spread of online education, there is a demand for improving the quality and effectiveness of educational content. However, traditional educational content has difficulty adapting to different learning styles and languages, which can result in reduced learning effectiveness. In particular, providing content in multiple languages ​​and at an individual learning pace is a challenge.

[0150] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0151] In this invention, the server includes means for receiving uploaded audio and video data, audio conversion means for converting audio data into other audio data, and video conversion means for converting video data into other video data, thereby enabling the audio and video data to be converted individually and personalized educational content to be generated and distributed.

[0152] "Means for receiving uploaded audio and video data" means the interface and associated software functionality through which a user provides audio and video data to the system.

[0153] "A speech conversion means for converting speech data into other speech data" refers to the deep learning models and associated processing algorithms used to convert the original speech data into a predetermined speech characteristic or language.

[0154] "Video transformation means for transforming video data into other video data" refers to the deep learning models and associated processing algorithms used to transform specific elements (e.g., facial features) in the original video into predetermined video characteristics or other video elements.

[0155] "Means for synthesizing and outputting converted audio and video data" refers to technology for combining converted audio data and video data to generate a single integrated file or stream and providing it to the user.

[0156] "Means for individually converting and generating personalized educational content" refers to a series of processes for customizing audio and video data for individual learners and creating optimized educational content.

[0157] "Means for delivering educational content" refers to the network infrastructure and related software capabilities required to deliver the generated educational content to students.

[0158] This invention relates to a system for generating and delivering personalized educational content in online education settings, which receives audio and video data uploaded by users, converts the audio and video data into other audio and video data, and synthesizes them to generate optimized educational content.

[0159] The server plays a central role in this system and utilizes the following hardware and software:

[0160] Hardware:

[0161] High-performance processor

[0162] Large capacity memory

[0163] Storage Devices

[0164] Network Interface

[0165] software:

[0166] Operating system (e.g. Linux)

[0167] Deep learning frameworks (e.g., PyTorch)

[0168] Video and audio processing libraries (e.g. FFmpeg, torchaudio)

[0169] Database management system (e.g. MySQL)

[0170] Data processing and calculation:

[0171] 1. Uploading and receiving data:

[0172] Through a web interface, users upload audio and video data for a presentation to a server, which has the means to receive the data and store the audio and video data as separate files.

[0173] 2. Audio data conversion:

[0174] The server uses deep learning models to convert voice data into other voice data, including processes that change voice characteristics and language based on user requests. The deep learning models used include "Speech2TextProcessor" and "Speech2TextForConditionalGeneration," both of which run on PyTorch.

[0175] 3. Video data conversion:

[0176] For video data, a deep learning model is used to transform specific elements (e.g., facial features) based on new facial images provided by the user. FFmpeg and PyTorch are used for video processing.

[0177] 4. Data synthesis and output:

[0178] The converted audio and video data is then recombined and output as new, optimized educational content, which is then distributed to students through educational institutions and online education platforms.

[0179] Examples:

[0180] For example, consider an educational institution that wants to upload audio and video of online lectures and convert them into personalized content in multiple languages. The educational institution's user uploads the original lecture data "lecture_original.mp4" to the system. The following prompt statements are used to specify audio and video conversion:

[0181] Example prompt sentence:

[0182] 1. Voice conversion: "Please translate this audio data into Japanese and resynthesize it in a calm voice."

[0183] 2. Video conversion: "Please replace the speaker's face in the lecture video with another face image."

[0184] The server processes these instructions and ultimately generates an optimized lecture video, "lecture_final.mp4," with the faces modified and translated into Japanese, which it then provides to the educational institution.

[0185] The above format will enable the creation and delivery of educational content suited to each student's learning style and language, which is expected to improve the quality and effectiveness of education.

[0186] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0187] Step 1:

[0188] Users upload audio and video data

[0189] Specific Operation: A user uploads an audio and video data file of a lecture or presentation (e.g., lecture_original.mp4) to the server using the web interface.

[0190] Input: Audio and video data files

[0191] Output: Audio and video data files stored on the server

[0192] Step 2:

[0193] The server receives the uploaded data and performs preprocessing.

[0194] Specific operation: The server splits the received audio and video data and saves them as audio data (e.g., extracted_audio.wav) and video data (e.g., video_no_audio.mp4). At the same time, it uses FFmpeg to perform preprocessing such as noise reduction and resolution conversion.

[0195] Input: Uploaded audio and video data files

[0196] Output: Preprocessed audio and video data files

[0197] Step 3:

[0198] The server converts the audio data

[0199] Specific operation: The server uses a speech conversion method (e.g., Speech2TextProcessor, Speech2TextForConditionalGeneration) to convert the voice data based on the user's prompts, and uses a generative AI model to modify the voice characteristics and language specified, generating a new voice data file (e.g., translated_audio.wav).

[0200] Input: Preprocessed audio data and speech-to-speech prompt

[0201] Output: Converted audio data file

[0202] Step 4:

[0203] The server converts the video data

[0204] Specific operation: The server converts specific elements (e.g., facial features) in the video data using a video conversion method (e.g., Face Swap, etc.) and generates a new video data file (e.g., converted_video.mp4) based on the prompt text.

[0205] Input: Preprocessed video data and video transformation prompt

[0206] Output: Converted video data file

[0207] Step 5:

[0208] The server synthesizes the converted audio and video data.

[0209] Specific operation: The server synthesizes the converted audio data (e.g., translated_audio.wav) and video data (e.g., converted_video.mp4) to generate a new integrated educational content file (e.g., final_presentation.mp4), using FFmpeg to synchronize the audio and video.

[0210] Input: Converted audio and video data files

[0211] Output: Composite educational content file

[0212] Step 6:

[0213] Download or distribute user-generated educational content

[0214] Specific operation: The user downloads the new educational content file (e.g., final_presentation.mp4) generated on the system or distributes it through the educational platform, allowing students to view personalized content.

[0215] Input: Synthesized educational content file

[0216] Output: User download or distribution of content

[0217] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0218] The present invention relates to a system that combines means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, means for synthesizing and outputting the converted audio and video data, and an emotion engine that recognizes user emotions. The system of the present invention utilizes generative AI and the emotion engine to encourage honest expression of opinions and improve communication in the workplace.

[0219] Program processing

[0220] 1. User upload of data

[0221] Users upload presentation audio and video files to the system. This is done using a common web interface. The user selects the files and sends them to the server, where the data is ingested into the system.

[0222] 2. Data preprocessing by the server

[0223] The server receives the uploaded file and separates the audio and video data, saving them as separate files and performing pre-processing such as noise reduction and resolution adjustment if necessary. This step ensures a smooth conversion process.

[0224] 3. Emotional engine for recognizing user emotions

[0225] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's voice intonation, facial expressions, or biometric sensor data to identify the user's emotional state. This information is reflected in the subsequent audio and video conversion process.

[0226] 4. Data transformation by the server

[0227] Based on the results of the emotion engine, the server applies a voice conversion model to convert the original audio data into a target voice. In this case, it uses a deep learning model to convert it into the voice of the project manager's colleague. The converted result is saved as "converted_audio.mp3."

[0228] The server then applies the video conversion model to convert the original video data into a target video, which is then converted into the face of the project manager's colleague using a deep learning model. The converted video data is saved as "converted_video.mp4."

[0229] 5. Data synthesis and output by the server

[0230] The converted audio and video data are then combined by the server and output as a new presentation file, which is saved as "final_presentation.mp4" and provided to the user.

[0231] Specific examples

[0232] For example, imagine a project manager at a company receives an important presentation from an executive, saved in a file called "original_presentation.mp4."

[0233] 1. The project manager (user) uploads "original_presentation.mp4" to the system.

[0234] 2. The server receives this file and separates the audio and video into "original_audio.mp3" and "original_video.mp4" respectively.

[0235] 3. The server uses the emotion engine to recognize the user's emotion, which is reflected in the subsequent conversion process.

[0236] 4. Based on the recognized emotion, the server applies a voice conversion model to convert "original_audio.mp3" into the voice of the project manager's colleague. The converted result is saved as "converted_audio.mp3."

[0237] 5. Next, the server applies the video conversion model to convert "original_video.mp4" into a video showing the colleague's face. The converted video data is saved as "converted_video.mp4."

[0238] 6. Finally, the server combines "converted_audio.mp3" and "converted_video.mp4" to create a new file called "final_presentation.mp4."

[0239] 7. The project manager (user) downloads this "final_presentation.mp4" and watches the presentation with the faces and voices of their colleagues.

[0240] This allows project managers, unlike executives, to listen to presentations in a more relaxed state and more easily express their opinions frankly. This is expected to lead to improved communication and constructive exchange of opinions within the workplace. In this way, the system of the present invention reduces communication barriers caused by differences in position and provides an effective means for improving the work environment.

[0241] The processing flow will be explained below.

[0242] Step 1:

[0243] Users upload presentation audio and video files to the system using a common web interface where the user selects the files and sends them to the server.

[0244] Step 2:

[0245] The server receives the uploaded file and separates the audio and video data. Specifically, the server extracts the audio from the received video file and saves the audio data as "original_audio.mp3" and the video data as "original_video.mp4."

[0246] Step 3:

[0247] The server performs pre-processing such as noise reduction and normalization on audio data, and resolution and frame rate adjustment on video data to ensure the conversion process proceeds efficiently and accurately.

[0248] Step 4:

[0249] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the voice intonation, facial expressions, or biometric sensor data provided when the user accesses the system to identify the user's emotional state. This emotion information is used in the subsequent conversion process.

[0250] Step 5:

[0251] Based on the emotional information recognized by the emotion engine, the server applies a voice conversion model to convert the original voice data into a target voice, for example, the voice of a colleague of the project manager, and the converted voice data is saved as "converted_audio.mp3."

[0252] Step 6:

[0253] Similarly, the server applies the video conversion model based on the emotion information recognized by the emotion engine to convert the original video data into a target video, specifically, the face of the project manager's colleague, and saves the converted video data as "converted_video.mp4."

[0254] Step 7:

[0255] The server combines the converted audio and video data to create a new presentation file. Specifically, it merges "converted_audio.mp3" and "converted_video.mp4" to create a new file called "final_presentation.mp4."

[0256] Step 8:

[0257] Users download "final_presentation.mp4" from the system. They can play this file and watch the presentation being given by their colleagues, with their faces and voices. This process creates an environment where participants can express their opinions frankly without feeling out of place when receiving a presentation from a senior executive.

[0258] Through these steps, the system of the present invention can provide users with an effective means of reducing status differences and improving communication within the workplace.

[0259] Example 2

[0260] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0261] Conventional presentation systems have the problem that it is difficult to exchange opinions frankly due to differences in hierarchical relationships and positions. In particular, constructive exchange of opinions and communication within the workplace is often hindered because the content of the presentation cannot be received in a more relaxed environment. This causes problems such as reduced efficiency and productivity throughout the organization.

[0262] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0263] In this invention, the server includes means for receiving uploaded audio and video data, audio conversion means for converting audio data into other audio data, video conversion means for converting video data into other video data, means for synthesizing and outputting the converted audio and video data, emotion recognition means for recognizing a user's emotion, means for reflecting the emotion information obtained by the emotion recognition means in the audio and video conversion process, and process control means for managing a series of processes from uploading to output. This allows users to receive presentations in different emotional states and facilitates frank exchange of opinions in a relaxed environment.

[0264] "Audio data" means data that records sound and is stored in digital or analog format.

[0265] "Video data" refers to data that records video and is stored in digital or analog format.

[0266] "Means for receiving" refers to devices or software that have the function of acquiring data and incorporating it into the system.

[0267] "Audio conversion means" refers to a device or software that has the function of converting original audio data into another audio data.

[0268] "Video conversion means" refers to a device or software that has the function of converting original video data into other video data.

[0269] "Means for synthesis and output" refers to devices or software that have the function of combining multiple pieces of data (audio, video, etc.) into the final output format.

[0270] "Emotion recognition means" refers to a device or software that has the function of analyzing audio, video, or biometric sensor data to identify a user's emotions.

[0271] "Process control means" refers to devices or software that have the function of managing the entire process in an orderly manner and properly executing each processing step.

[0272] A "machine learning model" refers to an algorithm or technique that learns patterns based on specific data and processes new data according to those patterns.

[0273] The system of the present invention converts uploaded audio and video data into other audio and video data, and finally synthesizes and outputs them. This system has the function of recognizing the user's emotions and reflecting this information in the audio and video conversion process.

[0274] Hardware and Software Configuration

[0275] The system implementation uses the following major hardware and software:

[0276] 1. Server: A central processing unit that receives, preprocesses, converts, synthesizes, and outputs data. Specifically, a server machine with high processing power is used.

[0277] 2. Web Interface: This is an interface via a web browser for users to upload audio and video data. Users access the file upload page in their browser, click the file selection button to select a file, and then press the upload button to send the data to the server.

[0278] 3. Audio analysis software: A library for processing the uploaded audio data, such as FFmpeg or OpenSmile, which separates, preprocesses, and converts the audio data.

[0279] 4. Video analysis software: Libraries for processing the uploaded video data, such as FFmpeg and OpenCV, which separate, preprocess, and convert the video data.

[0280] 5. Emotion Recognition Engine: Software for analyzing a user's emotions from audio and video data. It identifies their emotional state by analyzing their voice intonation and facial expressions.

[0281] 6. Machine learning models: These are models that use deep learning to convert audio and video data. Specifically, models such as WaveNet and FaceSwap are used.

[0282] Example of a system

[0283] For example, imagine a project manager at a company receives an important presentation, which is saved in a file called "original_presentation.mp4."

[0284] 1. The project manager (user) uploads "original_presentation.mp4" to the system. Using the web interface, select the file on the file upload page and submit it.

[0285] 2. The server receives this file and separates the audio and video data using FFmpeg to create audio data "original_audio.mp3" and video data "original_video.mp4".

[0286] 3. The server uses an emotion recognition engine to recognize the user's emotions. It uses OpenSmile and OpenCV to analyze the emotional state of the voice and video, and obtains information such as whether the user is excited or happy.

[0287] 4. The server applies a voice conversion model based on the results of the emotion recognition engine to convert "original_audio.mp3" into the colleague's voice. This conversion uses WaveNet, and the result is saved as "converted_audio.mp3."

[0288] 5. The server applies the video transformation model to convert "original_video.mp4" to the colleague's face. This transformation uses FaceSwap, and the result is saved as "converted_video.mp4."

[0289] 6. The server combines the converted audio and video data and saves it as the final file "final_presentation.mp4." It uses FFmpeg to combine the audio and video to generate a new multimedia file.

[0290] In this way, project managers can see and hear their colleagues' presentations, evaluate and exchange opinions in a relaxed environment.

[0291] Prompt Sentence Examples

[0292] "Please describe a system that uploads presentation audio and video data, uses an emotion recognition engine to convert the audio and video to that of a colleague, and finally generates a composite presentation file."

[0293] This system will enable users to exchange opinions frankly without worrying about hierarchical relationships or differences in position, and is expected to improve communication in the workplace.

[0294] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0295] Step 1:

[0296] The user uploads the presentation audio and video file ("original_presentation.mp4") to the system through a web interface: the user accesses the file upload page in their browser, clicks the file selection button to select the presentation file, and presses the 'Upload' button, which sends the selected file to the server.

[0297] Input: Presentation file "original_presentation.mp4"

[0298] Output: File upload request to the server

[0299] Step 2:

[0300] The server receives the uploaded "original_presentation.mp4" and separates the audio and video data into "original_audio.mp3" and "original_video.mp4" using the FFmpeg library.

[0301] Input: Uploaded presentation file "original_presentation.mp4"

[0302] Output: Separated audio data "original_audio.mp3" and video data "original_video.mp4"

[0303] Step 3:

[0304] The server preprocesses the received audio data "original_audio.mp3" and video data "original_video.mp4" by applying noise reduction to the audio data and adjusting the resolution of the video data.

[0305] Input: Separated audio data "original_audio.mp3", video data "original_video.mp4"

[0306] Output: Preprocessed audio and video data

[0307] Step 4:

[0308] The server uses the pre-processed audio and video data to recognize the user's emotions. It uses an emotion recognition engine to analyze intonation from the audio data and facial expressions from the video data to identify the user's emotional state.

[0309] Input: Preprocessed audio data, preprocessed video data

[0310] Output: User's emotional state data (e.g., excitement, joy)

[0311] Step 5:

[0312] Based on the recognized emotional state, the server applies a voice conversion model to convert the original audio data "original_audio.mp3" into the target audio data "converted_audio.mp3", using a deep learning model such as WaveNet.

[0313] Input: Preprocessed audio data "original_audio.mp3", user emotional state data

[0314] Output: Converted audio data "converted_audio.mp3"

[0315] Step 6:

[0316] Based on the recognized emotional state, the server applies a video conversion model to convert the original video data "original_video.mp4" into the target video data "converted_video.mp4." This conversion uses a deep learning model such as FaceSwap.

[0317] Input: Preprocessed video data "original_video.mp4", user emotional state data

[0318] Output: Converted video data "converted_video.mp4"

[0319] Step 7:

[0320] The server combines the converted audio data (converted_audio.mp3) and video data (converted_video.mp4) and outputs the combined data as a new presentation file (final_presentation.mp4). FFmpeg is used to combine the audio and video to generate a new multimedia file.

[0321] Input: Converted audio data "converted_audio.mp3", converted video data "converted_video.mp4"

[0322] Output: The final composite presentation file "final_presentation.mp4"

[0323] This series of processes allows users to watch presentations delivered in the presence and voice of their colleagues, and to evaluate and exchange opinions in a relaxed environment.

[0324] (Application example 2)

[0325] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0326] Conventional presentation systems have difficulty adapting the audio and video of a presentation to a specific target audience, and are unable to recognize and adapt the content based on user emotions, limiting their effectiveness in attracting audience attention. Furthermore, they lack the means to allow users to relax and express their honest opinions. This invention solves these problems and provides a more effective presentation adaptation and viewing experience.

[0327] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0328] In this invention, the server includes means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, means for synthesizing and outputting the converted audio and video data, emotion recognition means for recognizing a user's emotion, control means for converting the audio and video data based on the emotion information recognized by the emotion recognition means, and means for synthesizing and outputting the converted audio and video data in a viewable format. This allows users to convert uploaded presentations to suit specific targets and make adjustments based on emotions to create more effective and attractive presentations. This also allows viewers to watch the content in a relaxed atmosphere and more easily express their honest opinions.

[0329] "Upload" refers to the operation in which a user sends a local file to the server and imports the data into the system.

[0330] "Presentation audio data" is digital data containing audio information used in a presentation.

[0331] "Presentation video data" is digital data containing video information used in a presentation.

[0332] "Audio conversion means" refers to a function or device for converting uploaded presentation audio data into other audio data.

[0333] "Video conversion means" refers to a function or device for converting uploaded presentation video data into other video data.

[0334] "Emotion recognition means" refers to functions or devices for analyzing and recognizing emotions from the user's voice or video.

[0335] The "control means" refers to a function or device that supervises and controls the conversion of audio and video data based on the emotional information obtained by the emotional recognition means.

[0336] "Synthesis" refers to the process of combining multiple pieces of data into a single format.

[0337] A "deep learning model" is an artificial intelligence algorithm that uses multi-layered neural networks to learn from data and perform specific tasks.

[0338] A "prompt" is a document that contains instructions or questions given to a generative AI model, and is an input document used to obtain specific results.

[0339] The system required to implement this invention includes the following hardware and software. The main hardware required is a smartphone and a server. The software requires a web interface, deep learning models (e.g., TensorFlow, PyTorch), audio and video processing libraries (e.g., WaveNet, GANs), and emotion recognition APIs (e.g., Google Cloud Vision API, Microsoft Azure Emotion API).

[0340] The embodiment of this system mainly operates through the following process. The process begins with the user uploading presentation audio and video data. The server then separates the data and saves them as separate files. The server then uses emotion recognition means to analyze emotions from the user's audio and video. The obtained emotion information is reflected in the audio and video data conversion process.

[0341] As a concrete example, consider the following scenario: A user wants to convert an important internal presentation into a character played by a famous actor. The user uploads the audio and video data of the presentation from their smartphone to the application. This data is sent to the server, where it is separated into audio and video. Next, an emotion recognition API is used to analyze the user's emotions, and a deep learning model is used to convert the audio and video data into that of the target character. This conversion is performed using prompt sentences to provide specific instructions.

[0342] For example, you can use the following prompt:

[0343] "Translate this presentation into the audio and visuals of a famous movie actor, maintaining his signature friendly and calming tone."

[0344] The converted audio and video data is then recombined on the server side and output as a new presentation. Users can then easily view and share the resulting presentation on their smartphones. Through this entire process, users can view the presentation in a more relaxed state and more easily express their honest opinions.

[0345] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0346] Step 1:

[0347] Uploading data

[0348] The user uploads the presentation audio and video data to the application from their smartphone. The user selects the file and initiates the upload using the web interface. The input is the presentation file, and the output is the audio and video data file stored on the server. Specifically, the user selects the target file from the smartphone's file system and taps the send button, which uploads the file to the server.

[0349] Step 2:

[0350] Data Preprocessing

[0351] The server receives the uploaded presentation file and separates the audio and video data. The input is the uploaded presentation file (e.g., .mp4 or .pptx), and the output is separate audio (.mp3) and video (.mp4) files. During this process, preprocessing such as noise removal and resolution adjustment is performed. Specifically, the server analyzes the file and uses the appropriate processing library to separate and save the audio and video.

[0352] Step 3:

[0353] emotion recognition

[0354] The server uses an emotion recognition API to recognize the user's emotions from audio and video data. The input is the separated audio and video files and a prompt from the user, and the output is the recognized emotion data. Specifically, it analyzes the intonation of the voice and the facial features of the video, and obtains an emotion label using the emotion recognition API.

[0355] Step 4:

[0356] Audio data conversion

[0357] The server converts the voice data using a deep learning model based on the emotion recognition results and the prompt. The input is the original voice data, the prompt, and the recognized emotion data, and the output is a converted voice file (converted_audio.mp3) with the target voice characteristics. Specifically, the voice is processed using a voice conversion model such as WaveNet.

[0358] Step 5:

[0359] Video data conversion

[0360] The server converts the video data using a deep learning model based on the emotion recognition results and prompt text. The input is the original video data, the prompt text, and the recognized emotion data, and the output is a converted video file (converted_video.mp4) with the target video characteristics. Specifically, it uses GANs and other tools to process and change faces and specific parts of the video.

[0361] Step 6:

[0362] Data synthesis and output

[0363] The server synthesizes the converted audio and video data and generates a new presentation file (final_presentation.mp4). The input is the converted audio and video files, and the output is the final synthesized presentation file. Specifically, the audio and video are synchronized and synthesized using a multimedia processing library such as FFmpeg.

[0364] Step 7:

[0365] User Viewing and Sharing

[0366] The user downloads the generated presentation to their smartphone, where it can be viewed and shared. The input is the final presentation file, and the output is the user's viewing experience and sharing via social media. Specific actions include using the application to download the video and tapping the play button. The user can also use the share button to share via email or social media.

[0367] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0368] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0369] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0370] [Second embodiment]

[0371] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0372] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0373] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0374] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0375] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0376] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0377] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0378] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0379] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0380] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0381] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0382] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0383] The present invention relates to a system including means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, and means for synthesizing and outputting the converted audio and video data. The system of the present invention utilizes generative AI to encourage candid opinions and improve communication in the workplace.

[0384] Program processing

[0385] 1. User upload of data

[0386] Users upload the presentation's audio and video files to the system, which is done using a common web interface.

[0387] 2. Data preprocessing by the server

[0388] The server receives the uploaded file and separates the audio and video data. The audio and video are saved as separate files, and preprocessing such as noise reduction and resolution conversion is performed as needed.

[0389] 3. Data transformation by the server

[0390] The server uses deep learning models to convert audio and video data: the audio conversion model converts the original audio data into other audio data, and the video conversion model converts the original video data into other video data.

[0391] 4. Data synthesis and output by the server

[0392] The converted audio and video data are then combined by the server and output as a new presentation file, which is then provided to the user.

[0393] Specific examples

[0394] For example, imagine a project manager at a company receives an important presentation from an executive, saved in a file called "original_presentation.mp4."

[0395] 1. The project manager (user) uploads "original_presentation.mp4" to the system.

[0396] 2. The server receives this file and separates the audio and video into "original_audio.mp3" and "original_video.mp4" respectively.

[0397] 3. The server applies the voice conversion model to convert "original_audio.mp3" into the voice of the project manager's colleague, resulting in the audio file "converted_audio.mp3."

[0398] 4. The server then applies the video conversion model to convert "original_video.mp4" into a video showing the colleague's face, resulting in a video file called "converted_video.mp4."

[0399] 5. Finally, the server combines "converted_audio.mp3" and "converted_video.mp4" to create a new file called "final_presentation.mp4."

[0400] 6. The project manager (user) downloads this "final_presentation.mp4" and watches the presentation with the faces and voices of their colleagues.

[0401] This allows project managers, unlike executives, to listen to presentations in a more relaxed state and more easily express their opinions frankly. This is expected to lead to improved communication and constructive exchange of opinions within the workplace. In this way, the system of the present invention reduces communication barriers caused by differences in position and provides an effective means for improving the work environment.

[0402] The processing flow will be explained below.

[0403] Step 1:

[0404] Users upload presentation audio and video files to the system. This process is done using a web interface, where users select files and send them to the server, where they are received and stored.

[0405] Step 2:

[0406] The server receives the uploaded file and separates the audio and video data. The server then reads the received video file, extracts the audio, and saves it as a separate file. Specifically, the audio is saved as "original_audio.mp3" and the video is saved as "original_video.mp4."

[0407] Step 3:

[0408] The server pre-processes the audio data by denoising and normalising it, and also adjusts the resolution and frame rate of the video data as needed, preparing it for a more efficient and accurate conversion process.

[0409] Step 4:

[0410] The server uses a voice conversion model to convert the original audio data into a target voice. In this case, it uses a deep learning model to convert it into the voice of the project manager's colleague. The converted result is saved as "converted_audio.mp3."

[0411] Step 5:

[0412] The server uses a video conversion model to convert the original video data into a target video, in this case the face of the project manager's colleague, using a deep learning model. The converted video data is saved as "converted_video.mp4."

[0413] Step 6:

[0414] The server combines the converted audio and video data. Specifically, it merges "converted_audio.mp3" and "converted_video.mp4" to generate a new presentation file. This new file is saved as "final_presentation.mp4."

[0415] Step 7:

[0416] Users download "final_presentation.mp4" from the system. They can play this file and watch the presentation being given by their colleagues, with their faces and voices. This process creates an environment where participants feel more comfortable listening to a presentation from a senior executive and can express their opinions more frankly.

[0417] Through these steps, the system of the present invention can provide users with an effective means of reducing status differences and improving communication within the workplace.

[0418] Example 1

[0419] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0420] Current presentations rely on the presenter's specific audio and video, making it difficult for viewers to evaluate the content unbiased. Furthermore, when it comes to smooth communication in the workplace, differences in perspective can make it difficult for people to express their opinions frankly. Conventional systems are limited in the technology they use, making it difficult to efficiently convert and synthesize audio and video.

[0421] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0422] In this invention, the server includes means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, means for converting the audio data and video data using a deep learning model, means for performing noise reduction and resolution conversion as preprocessing, and means for synthesizing and outputting the converted audio and video data. This reduces communication barriers due to differences in positions in the workplace and enables constructive exchange of opinions.

[0423] "Upload" is the process by which a user submits data to the system for storage.

[0424] "Presentation audio data" is digital data of the audio used during a presentation.

[0425] "Presentation video data" refers to digital data of the video used during a presentation.

[0426] "Audio conversion means" refers to a method or device for converting original audio data into audio data of a different format or content.

[0427] "Video conversion means" refers to a method or device for converting original video data into video data of a different format or content.

[0428] A "deep learning model" is a machine learning model that uses a multi-layer neural network to convert and analyze data.

[0429] "Preprocessing" refers to the process of performing basic processing such as noise removal and resolution conversion before converting the data.

[0430] "Denoising" is the process of removing unwanted noise from audio data.

[0431] "Resolution conversion" is the process of changing the resolution of video data.

[0432] "Synthesis" is the process of combining multiple pieces of data (audio data or video data) into one piece of data.

[0433] "Output" refers to providing processed data in a form usable by a user.

[0434] The present invention is a system that receives audio and video data of a presentation, converts it, and outputs it. The system processes presentation data uploaded by a user and converts the audio and video data into other formats using a deep learning model, enabling more effective communication.

[0435] First, a user uploads the presentation audio and video files to the system using the web interface. For example, a user may upload a file called "original_presentation.mp4." To do this, the user accesses the specified URL through a web browser and clicks the upload button.

[0436] Next, the server receives the uploaded file and separates the audio and video data. Specifically, it uses a tool such as FFmpeg to save the audio and video as separate files. It then applies a noise reduction filter to the separated audio data and performs resolution conversion on the video data. For example, the audio data is saved as "original_audio.mp3" and the video data as "original_video.mp4."

[0437] The server then uses a deep learning model to convert the audio and video data. It uses an audio conversion model (for example, a model using TensorFlow or PyTorch) to convert the audio data and generate a new audio file called "converted_audio.mp3". Similarly, it uses a video conversion model to generate a new video file called "converted_video.mp4".

[0438] The converted audio and video data are then combined by the server and output as the final presentation file, "final_presentation.mp4." Using a media processing library such as FFmpeg, the audio and video data are combined into a single file. Users can then download this new file and view it in their own environment.

[0439] As a concrete example, consider the case where a project manager at a company receives an important presentation from an executive. This presentation is saved in "original_presentation.mp4." The project manager (user) uploads this file to the system. The server receives this file and separates the audio and video. After noise removal and resolution conversion, it applies audio conversion models and video conversion models to generate new data for each. Finally, it is provided to the user as "final_presentation.mp4."

[0440] Example prompt sentence:

[0441] "Please upload your presentation file. Filename: original_presentation.mp4"

[0442] "Separating audio and video data. Please wait a moment."

[0443] "Converting audio data. Progress: 40%"

[0444] "Converting video data. Progress: 70%"

[0445] "Synthesizing converted data. Progress: 90%"

[0446] "Conversion complete. Download your file. Filename: final_presentation.mp4"

[0447] In this way, the system of the present invention uses deep learning technology and various media processing tools to effectively convert and synthesize presentation data uploaded by users, thereby improving communication within the workplace.

[0448] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0449] Step 1: User uploads data

[0450] A user uploads a presentation's audio and video files to the system using a web interface. The input is a file called "original_presentation.mp4," which is sent to the server. Specifically, the user opens a web browser, accesses the specified URL, clicks the "Choose File" button, selects the desired file, and presses the "Upload" button.

[0451] input:

[0452] Upload file: original_presentation.mp4

[0453] output:

[0454] The file is sent to the server

[0455] Step 2: Server receives and separates data

[0456] The server receives the uploaded "original_presentation.mp4" and uses FFmpeg to separate the audio and video data. Specifically, it saves the received file in a specific directory and executes the FFmpeg command to generate the audio data "original_audio.mp3" and the video data "original_video.mp4."

[0457] input:

[0458] Upload file: original_presentation.mp4

[0459] output:

[0460] Separated audio data: original_audio.mp3

[0461] Separated video data: original_video.mp4

[0462] Step 3: Preprocessing the data on the server

[0463] The server performs preprocessing on the separated audio and video data, applying an audio filter to remove noise and converting the resolution. Specifically, it uses the FFmpeg command to remove noise from the audio data (e.g., ffmpeg -i original_audio.mp3 -af "highpass=f=200, lowpass=f=3000" cleaned_audio.mp3) and converts the resolution of the video data (e.g., ffmpeg -i original_video.mp4 -vf scale=1280:720 resized_video.mp4).

[0464] input:

[0465] Audio data: original_audio.mp3

[0466] Video data: original_video.mp4

[0467] output:

[0468] Noise-removed audio data: cleaned_audio.mp3

[0469] Resolution converted video data: resized_video.mp4

[0470] Step 4: Data transformation by the server

[0471] The server converts audio and video data using a deep learning model. Specifically, it inputs audio data into an audio conversion model using TensorFlow or PyTorch, generating a converted audio file called "converted_audio.mp3." Similarly, it inputs video data into a video conversion model and generates a converted video file called "converted_video.mp4" (e.g., python audio_conversion.py --input cleaned_audio.mp3 --output converted_audio.mp3 and python video_conversion.py --input resized_video.mp4 --output converted_video.mp4).

[0472] input:

[0473] Preprocessed audio data: cleaned_audio.mp3

[0474] Preprocessed video data: resized_video.mp4

[0475] output:

[0476] Converted audio data: converted_audio.mp3

[0477] Converted video data: converted_video.mp4

[0478] Step 5: Data synthesis by the server

[0479] The server combines the converted audio and video data to generate the final presentation file, "final_presentation.mp4." Specifically, it uses FFmpeg to execute a command to combine the audio and video into a single file (e.g., ffmpeg -i converted_video.mp4 -i converted_audio.mp3 -c:v copy -c:a aac final_presentation.mp4).

[0480] input:

[0481] Converted audio data: converted_audio.mp3

[0482] Converted video data: converted_video.mp4

[0483] output:

[0484] Final presentation file: final_presentation.mp4

[0485] Step 6: User downloads final file

[0486] The user downloads the generated final presentation file, “final_presentation.mp4,” via the web interface by clicking the download link provided by the system.

[0487] input:

[0488] Final presentation file: final_presentation.mp4

[0489] output:

[0490] Downloaded file: final_presentation.mp4

[0491] (Application example 1)

[0492] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0493] With the spread of online education, there is a demand for improving the quality and effectiveness of educational content. However, traditional educational content has difficulty adapting to different learning styles and languages, which can result in reduced learning effectiveness. In particular, providing content in multiple languages ​​and at an individual learning pace is a challenge.

[0494] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0495] In this invention, the server includes means for receiving uploaded audio and video data, audio conversion means for converting audio data into other audio data, and video conversion means for converting video data into other video data, thereby enabling the audio and video data to be converted individually and personalized educational content to be generated and distributed.

[0496] "Means for receiving uploaded audio and video data" means the interface and associated software functionality through which a user provides audio and video data to the system.

[0497] "A speech conversion means for converting speech data into other speech data" refers to the deep learning models and associated processing algorithms used to convert the original speech data into a predetermined speech characteristic or language.

[0498] "Video transformation means for transforming video data into other video data" refers to the deep learning models and associated processing algorithms used to transform specific elements (e.g., facial features) in the original video into predetermined video characteristics or other video elements.

[0499] "Means for synthesizing and outputting converted audio and video data" refers to technology for combining converted audio data and video data to generate a single integrated file or stream and providing it to the user.

[0500] "Means for individually converting and generating personalized educational content" refers to a series of processes for customizing audio and video data for individual learners and creating optimized educational content.

[0501] "Means for delivering educational content" refers to the network infrastructure and related software capabilities required to deliver the generated educational content to students.

[0502] This invention relates to a system for generating and delivering personalized educational content in online education settings, which receives audio and video data uploaded by users, converts the audio and video data into other audio and video data, and synthesizes them to generate optimized educational content.

[0503] The server plays a central role in this system and utilizes the following hardware and software:

[0504] Hardware:

[0505] High-performance processor

[0506] Large capacity memory

[0507] Storage Devices

[0508] Network Interface

[0509] software:

[0510] Operating system (e.g. Linux)

[0511] Deep learning frameworks (e.g., PyTorch)

[0512] Video and audio processing libraries (e.g. FFmpeg, torchaudio)

[0513] Database management system (e.g. MySQL)

[0514] Data processing and calculation:

[0515] 1. Uploading and receiving data:

[0516] Through a web interface, users upload audio and video data for a presentation to a server, which has the means to receive the data and store the audio and video data as separate files.

[0517] 2. Audio data conversion:

[0518] The server uses deep learning models to convert voice data into other voice data, including processes that change voice characteristics and language based on user requests. The deep learning models used include "Speech2TextProcessor" and "Speech2TextForConditionalGeneration," both of which run on PyTorch.

[0519] 3. Video data conversion:

[0520] For video data, a deep learning model is used to transform specific elements (e.g., facial features) based on new facial images provided by the user. FFmpeg and PyTorch are used for video processing.

[0521] 4. Data synthesis and output:

[0522] The converted audio and video data is then recombined and output as new, optimized educational content, which is then distributed to students through educational institutions and online education platforms.

[0523] Examples:

[0524] For example, consider an educational institution that wants to upload audio and video of online lectures and convert them into personalized content in multiple languages. The educational institution's user uploads the original lecture data "lecture_original.mp4" to the system. The following prompt statements are used to specify audio and video conversion:

[0525] Example prompt sentence:

[0526] 1. Voice conversion: "Please translate this audio data into Japanese and resynthesize it in a calm voice."

[0527] 2. Video conversion: "Please replace the speaker's face in the lecture video with another face image."

[0528] The server processes these instructions and ultimately generates an optimized lecture video, "lecture_final.mp4," with the faces modified and translated into Japanese, which it then provides to the educational institution.

[0529] The above format will enable the creation and delivery of educational content suited to each student's learning style and language, which is expected to improve the quality and effectiveness of education.

[0530] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0531] Step 1:

[0532] Users upload audio and video data

[0533] Specific Operation: A user uploads an audio and video data file of a lecture or presentation (e.g., lecture_original.mp4) to the server using the web interface.

[0534] Input: Audio and video data files

[0535] Output: Audio and video data files stored on the server

[0536] Step 2:

[0537] The server receives the uploaded data and performs preprocessing.

[0538] Specific operation: The server splits the received audio and video data and saves them as audio data (e.g., extracted_audio.wav) and video data (e.g., video_no_audio.mp4). At the same time, it uses FFmpeg to perform preprocessing such as noise reduction and resolution conversion.

[0539] Input: Uploaded audio and video data files

[0540] Output: Preprocessed audio and video data files

[0541] Step 3:

[0542] The server converts the audio data

[0543] Specific operation: The server uses a speech conversion method (e.g., Speech2TextProcessor, Speech2TextForConditionalGeneration) to convert the voice data based on the user's prompts, and uses a generative AI model to modify the voice characteristics and language specified, generating a new voice data file (e.g., translated_audio.wav).

[0544] Input: Preprocessed audio data and speech-to-speech prompt

[0545] Output: Converted audio data file

[0546] Step 4:

[0547] The server converts the video data

[0548] Specific operation: The server converts specific elements (e.g., facial features) in the video data using a video conversion method (e.g., Face Swap, etc.) and generates a new video data file (e.g., converted_video.mp4) based on the prompt text.

[0549] Input: Preprocessed video data and video transformation prompt

[0550] Output: Converted video data file

[0551] Step 5:

[0552] The server synthesizes the converted audio and video data.

[0553] Specific operation: The server synthesizes the converted audio data (e.g., translated_audio.wav) and video data (e.g., converted_video.mp4) to generate a new integrated educational content file (e.g., final_presentation.mp4), using FFmpeg to synchronize the audio and video.

[0554] Input: Converted audio and video data files

[0555] Output: Composite educational content file

[0556] Step 6:

[0557] Download or distribute user-generated educational content

[0558] Specific operation: The user downloads the new educational content file (e.g., final_presentation.mp4) generated on the system or distributes it through the educational platform, allowing students to view personalized content.

[0559] Input: Synthesized educational content file

[0560] Output: User download or distribution of content

[0561] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0562] The present invention relates to a system that combines means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, means for synthesizing and outputting the converted audio and video data, and an emotion engine that recognizes user emotions. The system of the present invention utilizes generative AI and the emotion engine to encourage honest expression of opinions and improve communication in the workplace.

[0563] Program processing

[0564] 1. User upload of data

[0565] Users upload presentation audio and video files to the system. This is done using a common web interface. The user selects the files and sends them to the server, where the data is ingested into the system.

[0566] 2. Data preprocessing by the server

[0567] The server receives the uploaded file and separates the audio and video data, saving them as separate files and performing pre-processing such as noise reduction and resolution adjustment if necessary. This step ensures a smooth conversion process.

[0568] 3. Emotional engine for recognizing user emotions

[0569] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's voice intonation, facial expressions, or biometric sensor data to identify the user's emotional state. This information is reflected in the subsequent audio and video conversion process.

[0570] 4. Data transformation by the server

[0571] Based on the results of the emotion engine, the server applies a voice conversion model to convert the original audio data into a target voice. In this case, it uses a deep learning model to convert it into the voice of the project manager's colleague. The converted result is saved as "converted_audio.mp3."

[0572] The server then applies the video conversion model to convert the original video data into a target video, which is then converted into the face of the project manager's colleague using a deep learning model. The converted video data is saved as "converted_video.mp4."

[0573] 5. Data synthesis and output by the server

[0574] The converted audio and video data are then combined by the server and output as a new presentation file, which is saved as "final_presentation.mp4" and provided to the user.

[0575] Specific examples

[0576] For example, imagine a project manager at a company receives an important presentation from an executive, saved in a file called "original_presentation.mp4."

[0577] 1. The project manager (user) uploads "original_presentation.mp4" to the system.

[0578] 2. The server receives this file and separates the audio and video into "original_audio.mp3" and "original_video.mp4" respectively.

[0579] 3. The server uses the emotion engine to recognize the user's emotion, which is reflected in the subsequent conversion process.

[0580] 4. Based on the recognized emotion, the server applies a voice conversion model to convert "original_audio.mp3" into the voice of the project manager's colleague. The converted result is saved as "converted_audio.mp3."

[0581] 5. Next, the server applies the video conversion model to convert "original_video.mp4" into a video showing the colleague's face. The converted video data is saved as "converted_video.mp4."

[0582] 6. Finally, the server combines "converted_audio.mp3" and "converted_video.mp4" to create a new file called "final_presentation.mp4."

[0583] 7. The project manager (user) downloads this "final_presentation.mp4" and watches the presentation with the faces and voices of their colleagues.

[0584] This allows project managers, unlike executives, to listen to presentations in a more relaxed state and more easily express their opinions frankly. This is expected to lead to improved communication and constructive exchange of opinions within the workplace. In this way, the system of the present invention reduces communication barriers caused by differences in position and provides an effective means for improving the work environment.

[0585] The processing flow will be explained below.

[0586] Step 1:

[0587] Users upload presentation audio and video files to the system using a common web interface where the user selects the files and sends them to the server.

[0588] Step 2:

[0589] The server receives the uploaded file and separates the audio and video data. Specifically, the server extracts the audio from the received video file and saves the audio data as "original_audio.mp3" and the video data as "original_video.mp4."

[0590] Step 3:

[0591] The server performs pre-processing such as noise reduction and normalization on audio data, and resolution and frame rate adjustment on video data to ensure the conversion process proceeds efficiently and accurately.

[0592] Step 4:

[0593] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the voice intonation, facial expressions, or biometric sensor data provided when the user accesses the system to identify the user's emotional state. This emotion information is used in the subsequent conversion process.

[0594] Step 5:

[0595] Based on the emotional information recognized by the emotion engine, the server applies a voice conversion model to convert the original voice data into a target voice, for example, the voice of a colleague of the project manager, and the converted voice data is saved as "converted_audio.mp3."

[0596] Step 6:

[0597] Similarly, the server applies the video conversion model based on the emotion information recognized by the emotion engine to convert the original video data into a target video, specifically, the face of the project manager's colleague, and saves the converted video data as "converted_video.mp4."

[0598] Step 7:

[0599] The server combines the converted audio and video data to create a new presentation file. Specifically, it merges "converted_audio.mp3" and "converted_video.mp4" to create a new file called "final_presentation.mp4."

[0600] Step 8:

[0601] Users download "final_presentation.mp4" from the system. They can play this file and watch the presentation being given by their colleagues, with their faces and voices. This process creates an environment where participants can express their opinions frankly without feeling out of place when receiving a presentation from a senior executive.

[0602] Through these steps, the system of the present invention can provide users with an effective means of reducing status differences and improving communication within the workplace.

[0603] Example 2

[0604] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0605] Conventional presentation systems have the problem that it is difficult to exchange opinions frankly due to differences in hierarchical relationships and positions. In particular, constructive exchange of opinions and communication within the workplace is often hindered because the content of the presentation cannot be received in a more relaxed environment. This causes problems such as reduced efficiency and productivity throughout the organization.

[0606] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0607] In this invention, the server includes means for receiving uploaded audio and video data, audio conversion means for converting audio data into other audio data, video conversion means for converting video data into other video data, means for synthesizing and outputting the converted audio and video data, emotion recognition means for recognizing a user's emotion, means for reflecting the emotion information obtained by the emotion recognition means in the audio and video conversion process, and process control means for managing a series of processes from uploading to output. This allows users to receive presentations in different emotional states and facilitates frank exchange of opinions in a relaxed environment.

[0608] "Audio data" means data that records sound and is stored in digital or analog format.

[0609] "Video data" refers to data that records video and is stored in digital or analog format.

[0610] "Means for receiving" refers to devices or software that have the function of acquiring data and incorporating it into the system.

[0611] "Audio conversion means" refers to a device or software that has the function of converting original audio data into another audio data.

[0612] "Video conversion means" refers to a device or software that has the function of converting original video data into other video data.

[0613] "Means for synthesis and output" refers to devices or software that have the function of combining multiple pieces of data (audio, video, etc.) into the final output format.

[0614] "Emotion recognition means" refers to a device or software that has the function of analyzing audio, video, or biometric sensor data to identify a user's emotions.

[0615] "Process control means" refers to devices or software that have the function of managing the entire process in an orderly manner and properly executing each processing step.

[0616] A "machine learning model" refers to an algorithm or technique that learns patterns based on specific data and processes new data according to those patterns.

[0617] The system of the present invention converts uploaded audio and video data into other audio and video data, and finally synthesizes and outputs them. This system has the function of recognizing the user's emotions and reflecting this information in the audio and video conversion process.

[0618] Hardware and Software Configuration

[0619] The system implementation uses the following major hardware and software:

[0620] 1. Server: A central processing unit that receives, preprocesses, converts, synthesizes, and outputs data. Specifically, a server machine with high processing power is used.

[0621] 2. Web Interface: This is an interface via a web browser for users to upload audio and video data. Users access the file upload page in their browser, click the file selection button to select a file, and then press the upload button to send the data to the server.

[0622] 3. Audio analysis software: A library for processing the uploaded audio data, such as FFmpeg or OpenSmile, which separates, preprocesses, and converts the audio data.

[0623] 4. Video analysis software: Libraries for processing the uploaded video data, such as FFmpeg and OpenCV, which separate, preprocess, and convert the video data.

[0624] 5. Emotion Recognition Engine: Software for analyzing a user's emotions from audio and video data. It identifies their emotional state by analyzing their voice intonation and facial expressions.

[0625] 6. Machine learning models: These are models that use deep learning to convert audio and video data. Specifically, models such as WaveNet and FaceSwap are used.

[0626] Example of a system

[0627] For example, imagine a project manager at a company receives an important presentation, which is saved in a file called "original_presentation.mp4."

[0628] 1. The project manager (user) uploads "original_presentation.mp4" to the system. Using the web interface, select the file on the file upload page and submit it.

[0629] 2. The server receives this file and separates the audio and video data using FFmpeg to create audio data "original_audio.mp3" and video data "original_video.mp4".

[0630] 3. The server uses an emotion recognition engine to recognize the user's emotions. It uses OpenSmile and OpenCV to analyze the emotional state of the voice and video, and obtains information such as whether the user is excited or happy.

[0631] 4. The server applies a voice conversion model based on the results of the emotion recognition engine to convert "original_audio.mp3" into the colleague's voice. This conversion uses WaveNet, and the result is saved as "converted_audio.mp3."

[0632] 5. The server applies the video transformation model to convert "original_video.mp4" to the colleague's face. This transformation uses FaceSwap, and the result is saved as "converted_video.mp4."

[0633] 6. The server combines the converted audio and video data and saves it as the final file "final_presentation.mp4." It uses FFmpeg to combine the audio and video to generate a new multimedia file.

[0634] In this way, project managers can see and hear their colleagues' presentations, evaluate and exchange opinions in a relaxed environment.

[0635] Prompt Sentence Examples

[0636] "Please describe a system that uploads presentation audio and video data, uses an emotion recognition engine to convert the audio and video to that of a colleague, and finally generates a composite presentation file."

[0637] This system will enable users to exchange opinions frankly without worrying about hierarchical relationships or differences in position, and is expected to improve communication in the workplace.

[0638] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0639] Step 1:

[0640] The user uploads the presentation audio and video file ("original_presentation.mp4") to the system through a web interface: the user accesses the file upload page in their browser, clicks the file selection button to select the presentation file, and presses the 'Upload' button, which sends the selected file to the server.

[0641] Input: Presentation file "original_presentation.mp4"

[0642] Output: File upload request to the server

[0643] Step 2:

[0644] The server receives the uploaded "original_presentation.mp4" and separates the audio and video data into "original_audio.mp3" and "original_video.mp4" using the FFmpeg library.

[0645] Input: Uploaded presentation file "original_presentation.mp4"

[0646] Output: Separated audio data "original_audio.mp3" and video data "original_video.mp4"

[0647] Step 3:

[0648] The server preprocesses the received audio data "original_audio.mp3" and video data "original_video.mp4" by applying noise reduction to the audio data and adjusting the resolution of the video data.

[0649] Input: Separated audio data "original_audio.mp3", video data "original_video.mp4"

[0650] Output: Preprocessed audio and video data

[0651] Step 4:

[0652] The server uses the pre-processed audio and video data to recognize the user's emotions. It uses an emotion recognition engine to analyze intonation from the audio data and facial expressions from the video data to identify the user's emotional state.

[0653] Input: Preprocessed audio data, preprocessed video data

[0654] Output: User's emotional state data (e.g., excitement, joy)

[0655] Step 5:

[0656] Based on the recognized emotional state, the server applies a voice conversion model to convert the original audio data "original_audio.mp3" into the target audio data "converted_audio.mp3", using a deep learning model such as WaveNet.

[0657] Input: Preprocessed audio data "original_audio.mp3", user emotional state data

[0658] Output: Converted audio data "converted_audio.mp3"

[0659] Step 6:

[0660] Based on the recognized emotional state, the server applies a video conversion model to convert the original video data "original_video.mp4" into the target video data "converted_video.mp4." This conversion uses a deep learning model such as FaceSwap.

[0661] Input: Preprocessed video data "original_video.mp4", user emotional state data

[0662] Output: Converted video data "converted_video.mp4"

[0663] Step 7:

[0664] The server combines the converted audio data (converted_audio.mp3) and video data (converted_video.mp4) and outputs the combined data as a new presentation file (final_presentation.mp4). FFmpeg is used to combine the audio and video to generate a new multimedia file.

[0665] Input: Converted audio data "converted_audio.mp3", converted video data "converted_video.mp4"

[0666] Output: The final composite presentation file "final_presentation.mp4"

[0667] This series of processes allows users to watch presentations delivered in the presence and voice of their colleagues, and to evaluate and exchange opinions in a relaxed environment.

[0668] (Application example 2)

[0669] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0670] Conventional presentation systems have difficulty adapting the audio and video of a presentation to a specific target audience, and are unable to recognize and adapt the content based on user emotions, limiting their effectiveness in attracting audience attention. Furthermore, they lack the means to allow users to relax and express their honest opinions. This invention solves these problems and provides a more effective presentation adaptation and viewing experience.

[0671] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0672] In this invention, the server includes means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, means for synthesizing and outputting the converted audio and video data, emotion recognition means for recognizing a user's emotion, control means for converting the audio and video data based on the emotion information recognized by the emotion recognition means, and means for synthesizing and outputting the converted audio and video data in a viewable format. This allows users to convert uploaded presentations to suit specific targets and make adjustments based on emotions to create more effective and attractive presentations. This also allows viewers to watch the content in a relaxed atmosphere and more easily express their honest opinions.

[0673] "Upload" refers to the operation in which a user sends a local file to the server and imports the data into the system.

[0674] "Presentation audio data" is digital data containing audio information used in a presentation.

[0675] "Presentation video data" is digital data containing video information used in a presentation.

[0676] "Audio conversion means" refers to a function or device for converting uploaded presentation audio data into other audio data.

[0677] "Video conversion means" refers to a function or device for converting uploaded presentation video data into other video data.

[0678] "Emotion recognition means" refers to functions or devices for analyzing and recognizing emotions from the user's voice or video.

[0679] The "control means" refers to a function or device that supervises and controls the conversion of audio and video data based on the emotional information obtained by the emotional recognition means.

[0680] "Synthesis" refers to the process of combining multiple pieces of data into a single format.

[0681] A "deep learning model" is an artificial intelligence algorithm that uses multi-layered neural networks to learn from data and perform specific tasks.

[0682] A "prompt" is a document that contains instructions or questions given to a generative AI model, and is an input document used to obtain specific results.

[0683] The system required to implement this invention includes the following hardware and software. The main hardware required is a smartphone and a server. The software requires a web interface, deep learning models (e.g., TensorFlow, PyTorch), audio and video processing libraries (e.g., WaveNet, GANs), and emotion recognition APIs (e.g., Google Cloud Vision API, Microsoft Azure Emotion API).

[0684] The embodiment of this system mainly operates through the following process. The process begins with the user uploading presentation audio and video data. The server then separates the data and saves them as separate files. The server then uses emotion recognition means to analyze emotions from the user's audio and video. The obtained emotion information is reflected in the audio and video data conversion process.

[0685] As a concrete example, consider the following scenario: A user wants to convert an important internal presentation into a character played by a famous actor. The user uploads the audio and video data of the presentation from their smartphone to the application. This data is sent to the server, where it is separated into audio and video. Next, an emotion recognition API is used to analyze the user's emotions, and a deep learning model is used to convert the audio and video data into that of the target character. This conversion is performed using prompt sentences to provide specific instructions.

[0686] For example, you can use the following prompt:

[0687] "Translate this presentation into the audio and visuals of a famous movie actor, maintaining his signature friendly and calming tone."

[0688] The converted audio and video data is then recombined on the server side and output as a new presentation. Users can then easily view and share the resulting presentation on their smartphones. Through this entire process, users can view the presentation in a more relaxed state and more easily express their honest opinions.

[0689] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0690] Step 1:

[0691] Uploading data

[0692] The user uploads the presentation audio and video data to the application from their smartphone. The user selects the file and initiates the upload using the web interface. The input is the presentation file, and the output is the audio and video data file stored on the server. Specifically, the user selects the target file from the smartphone's file system and taps the send button, which uploads the file to the server.

[0693] Step 2:

[0694] Data Preprocessing

[0695] The server receives the uploaded presentation file and separates the audio and video data. The input is the uploaded presentation file (e.g., .mp4 or .pptx), and the output is separate audio (.mp3) and video (.mp4) files. During this process, preprocessing such as noise removal and resolution adjustment is performed. Specifically, the server analyzes the file and uses the appropriate processing library to separate and save the audio and video.

[0696] Step 3:

[0697] emotion recognition

[0698] The server uses an emotion recognition API to recognize the user's emotions from audio and video data. The input is the separated audio and video files and a prompt from the user, and the output is the recognized emotion data. Specifically, it analyzes the intonation of the voice and the facial features of the video, and obtains an emotion label using the emotion recognition API.

[0699] Step 4:

[0700] Audio data conversion

[0701] The server converts the voice data using a deep learning model based on the emotion recognition results and the prompt. The input is the original voice data, the prompt, and the recognized emotion data, and the output is a converted voice file (converted_audio.mp3) with the target voice characteristics. Specifically, the voice is processed using a voice conversion model such as WaveNet.

[0702] Step 5:

[0703] Video data conversion

[0704] The server converts the video data using a deep learning model based on the emotion recognition results and prompt text. The input is the original video data, the prompt text, and the recognized emotion data, and the output is a converted video file (converted_video.mp4) with the target video characteristics. Specifically, it uses GANs and other tools to process and change faces and specific parts of the video.

[0705] Step 6:

[0706] Data synthesis and output

[0707] The server synthesizes the converted audio and video data and generates a new presentation file (final_presentation.mp4). The input is the converted audio and video files, and the output is the final synthesized presentation file. Specifically, the audio and video are synchronized and synthesized using a multimedia processing library such as FFmpeg.

[0708] Step 7:

[0709] User Viewing and Sharing

[0710] The user downloads the generated presentation to their smartphone, where it can be viewed and shared. The input is the final presentation file, and the output is the user's viewing experience and sharing via social media. Specific actions include using the application to download the video and tapping the play button. The user can also use the share button to share via email or social media.

[0711] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0712] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0713] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0714] [Third embodiment]

[0715] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0716] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0717] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0718] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0719] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0720] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0721] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0722] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0723] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0724] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0725] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0726] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0727] The present invention relates to a system including means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, and means for synthesizing and outputting the converted audio and video data. The system of the present invention utilizes generative AI to encourage candid opinions and improve communication in the workplace.

[0728] Program processing

[0729] 1. User upload of data

[0730] Users upload the presentation's audio and video files to the system, which is done using a common web interface.

[0731] 2. Data preprocessing by the server

[0732] The server receives the uploaded file and separates the audio and video data. The audio and video are saved as separate files, and preprocessing such as noise reduction and resolution conversion is performed as needed.

[0733] 3. Data transformation by the server

[0734] The server uses deep learning models to convert audio and video data: the audio conversion model converts the original audio data into other audio data, and the video conversion model converts the original video data into other video data.

[0735] 4. Data synthesis and output by the server

[0736] The converted audio and video data are then combined by the server and output as a new presentation file, which is then provided to the user.

[0737] Specific examples

[0738] For example, imagine a project manager at a company receives an important presentation from an executive, saved in a file called "original_presentation.mp4."

[0739] 1. The project manager (user) uploads "original_presentation.mp4" to the system.

[0740] 2. The server receives this file and separates the audio and video into "original_audio.mp3" and "original_video.mp4" respectively.

[0741] 3. The server applies the voice conversion model to convert "original_audio.mp3" into the voice of the project manager's colleague, resulting in the audio file "converted_audio.mp3."

[0742] 4. The server then applies the video conversion model to convert "original_video.mp4" into a video showing the colleague's face, resulting in a video file called "converted_video.mp4."

[0743] 5. Finally, the server combines "converted_audio.mp3" and "converted_video.mp4" to create a new file called "final_presentation.mp4."

[0744] 6. The project manager (user) downloads this "final_presentation.mp4" and watches the presentation with the faces and voices of their colleagues.

[0745] This allows project managers, unlike executives, to listen to presentations in a more relaxed state and more easily express their opinions frankly. This is expected to lead to improved communication and constructive exchange of opinions within the workplace. In this way, the system of the present invention reduces communication barriers caused by differences in position and provides an effective means for improving the work environment.

[0746] The processing flow will be explained below.

[0747] Step 1:

[0748] Users upload presentation audio and video files to the system. This process is done using a web interface, where users select files and send them to the server, where they are received and stored.

[0749] Step 2:

[0750] The server receives the uploaded file and separates the audio and video data. The server then reads the received video file, extracts the audio, and saves it as a separate file. Specifically, the audio is saved as "original_audio.mp3" and the video is saved as "original_video.mp4."

[0751] Step 3:

[0752] The server pre-processes the audio data by denoising and normalising it, and also adjusts the resolution and frame rate of the video data as needed, preparing it for a more efficient and accurate conversion process.

[0753] Step 4:

[0754] The server uses a voice conversion model to convert the original audio data into a target voice. In this case, it uses a deep learning model to convert it into the voice of the project manager's colleague. The converted result is saved as "converted_audio.mp3."

[0755] Step 5:

[0756] The server uses a video conversion model to convert the original video data into a target video, in this case the face of the project manager's colleague, using a deep learning model. The converted video data is saved as "converted_video.mp4."

[0757] Step 6:

[0758] The server combines the converted audio and video data. Specifically, it merges "converted_audio.mp3" and "converted_video.mp4" to generate a new presentation file. This new file is saved as "final_presentation.mp4."

[0759] Step 7:

[0760] Users download "final_presentation.mp4" from the system. They can play this file and watch the presentation being given by their colleagues, with their faces and voices. This process creates an environment where participants feel more comfortable listening to a presentation from a senior executive and can express their opinions more frankly.

[0761] Through these steps, the system of the present invention can provide users with an effective means of reducing status differences and improving communication within the workplace.

[0762] Example 1

[0763] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0764] Current presentations rely on the presenter's specific audio and video, making it difficult for viewers to evaluate the content unbiased. Furthermore, when it comes to smooth communication in the workplace, differences in perspective can make it difficult for people to express their opinions frankly. Conventional systems are limited in the technology they use, making it difficult to efficiently convert and synthesize audio and video.

[0765] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0766] In this invention, the server includes means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, means for converting the audio data and video data using a deep learning model, means for performing noise reduction and resolution conversion as preprocessing, and means for synthesizing and outputting the converted audio and video data. This reduces communication barriers due to differences in positions in the workplace and enables constructive exchange of opinions.

[0767] "Upload" is the process by which a user submits data to the system for storage.

[0768] "Presentation audio data" is digital data of the audio used during a presentation.

[0769] "Presentation video data" refers to digital data of the video used during a presentation.

[0770] "Audio conversion means" refers to a method or device for converting original audio data into audio data of a different format or content.

[0771] "Video conversion means" refers to a method or device for converting original video data into video data of a different format or content.

[0772] A "deep learning model" is a machine learning model that uses a multi-layer neural network to convert and analyze data.

[0773] "Preprocessing" refers to the process of performing basic processing such as noise removal and resolution conversion before converting the data.

[0774] "Denoising" is the process of removing unwanted noise from audio data.

[0775] "Resolution conversion" is the process of changing the resolution of video data.

[0776] "Synthesis" is the process of combining multiple pieces of data (audio data or video data) into one piece of data.

[0777] "Output" refers to providing processed data in a form usable by a user.

[0778] The present invention is a system that receives audio and video data of a presentation, converts it, and outputs it. The system processes presentation data uploaded by a user and converts the audio and video data into other formats using a deep learning model, enabling more effective communication.

[0779] First, a user uploads the presentation audio and video files to the system using the web interface. For example, a user may upload a file called "original_presentation.mp4." To do this, the user accesses the specified URL through a web browser and clicks the upload button.

[0780] Next, the server receives the uploaded file and separates the audio and video data. Specifically, it uses a tool such as FFmpeg to save the audio and video as separate files. It then applies a noise reduction filter to the separated audio data and performs resolution conversion on the video data. For example, the audio data is saved as "original_audio.mp3" and the video data as "original_video.mp4."

[0781] The server then uses a deep learning model to convert the audio and video data. It uses an audio conversion model (for example, a model using TensorFlow or PyTorch) to convert the audio data and generate a new audio file called "converted_audio.mp3". Similarly, it uses a video conversion model to generate a new video file called "converted_video.mp4".

[0782] The converted audio and video data are then combined by the server and output as the final presentation file, "final_presentation.mp4." Using a media processing library such as FFmpeg, the audio and video data are combined into a single file. Users can then download this new file and view it in their own environment.

[0783] As a concrete example, consider the case where a project manager at a company receives an important presentation from an executive. This presentation is saved in "original_presentation.mp4." The project manager (user) uploads this file to the system. The server receives this file and separates the audio and video. After noise removal and resolution conversion, it applies audio conversion models and video conversion models to generate new data for each. Finally, it is provided to the user as "final_presentation.mp4."

[0784] Example prompt sentence:

[0785] "Please upload your presentation file. Filename: original_presentation.mp4"

[0786] "Separating audio and video data. Please wait a moment."

[0787] "Converting audio data. Progress: 40%"

[0788] "Converting video data. Progress: 70%"

[0789] "Synthesizing converted data. Progress: 90%"

[0790] "Conversion complete. Download your file. Filename: final_presentation.mp4"

[0791] In this way, the system of the present invention uses deep learning technology and various media processing tools to effectively convert and synthesize presentation data uploaded by users, thereby improving communication within the workplace.

[0792] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0793] Step 1: User uploads data

[0794] A user uploads a presentation's audio and video files to the system using a web interface. The input is a file called "original_presentation.mp4," which is sent to the server. Specifically, the user opens a web browser, accesses the specified URL, clicks the "Choose File" button, selects the desired file, and presses the "Upload" button.

[0795] input:

[0796] Upload file: original_presentation.mp4

[0797] output:

[0798] The file is sent to the server

[0799] Step 2: Server receives and separates data

[0800] The server receives the uploaded "original_presentation.mp4" and uses FFmpeg to separate the audio and video data. Specifically, it saves the received file in a specific directory and executes the FFmpeg command to generate the audio data "original_audio.mp3" and the video data "original_video.mp4."

[0801] input:

[0802] Upload file: original_presentation.mp4

[0803] output:

[0804] Separated audio data: original_audio.mp3

[0805] Separated video data: original_video.mp4

[0806] Step 3: Preprocessing the data on the server

[0807] The server performs preprocessing on the separated audio and video data, applying an audio filter to remove noise and converting the resolution. Specifically, it uses the FFmpeg command to remove noise from the audio data (e.g., ffmpeg -i original_audio.mp3 -af "highpass=f=200, lowpass=f=3000" cleaned_audio.mp3) and converts the resolution of the video data (e.g., ffmpeg -i original_video.mp4 -vf scale=1280:720 resized_video.mp4).

[0808] input:

[0809] Audio data: original_audio.mp3

[0810] Video data: original_video.mp4

[0811] output:

[0812] Noise-removed audio data: cleaned_audio.mp3

[0813] Resolution converted video data: resized_video.mp4

[0814] Step 4: Data transformation by the server

[0815] The server converts audio and video data using a deep learning model. Specifically, it inputs audio data into an audio conversion model using TensorFlow or PyTorch, generating a converted audio file called "converted_audio.mp3." Similarly, it inputs video data into a video conversion model and generates a converted video file called "converted_video.mp4" (e.g., python audio_conversion.py --input cleaned_audio.mp3 --output converted_audio.mp3 and python video_conversion.py --input resized_video.mp4 --output converted_video.mp4).

[0816] input:

[0817] Preprocessed audio data: cleaned_audio.mp3

[0818] Preprocessed video data: resized_video.mp4

[0819] output:

[0820] Converted audio data: converted_audio.mp3

[0821] Converted video data: converted_video.mp4

[0822] Step 5: Data synthesis by the server

[0823] The server combines the converted audio and video data to generate the final presentation file, "final_presentation.mp4." Specifically, it uses FFmpeg to execute a command to combine the audio and video into a single file (e.g., ffmpeg -i converted_video.mp4 -i converted_audio.mp3 -c:v copy -c:a aac final_presentation.mp4).

[0824] input:

[0825] Converted audio data: converted_audio.mp3

[0826] Converted video data: converted_video.mp4

[0827] output:

[0828] Final presentation file: final_presentation.mp4

[0829] Step 6: User downloads final file

[0830] The user downloads the generated final presentation file, “final_presentation.mp4,” via the web interface by clicking the download link provided by the system.

[0831] input:

[0832] Final presentation file: final_presentation.mp4

[0833] output:

[0834] Downloaded file: final_presentation.mp4

[0835] (Application example 1)

[0836] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0837] With the spread of online education, there is a demand for improving the quality and effectiveness of educational content. However, traditional educational content has difficulty adapting to different learning styles and languages, which can result in reduced learning effectiveness. In particular, providing content in multiple languages ​​and at an individual learning pace is a challenge.

[0838] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0839] In this invention, the server includes means for receiving uploaded audio and video data, audio conversion means for converting audio data into other audio data, and video conversion means for converting video data into other video data, thereby enabling the audio and video data to be converted individually and personalized educational content to be generated and distributed.

[0840] "Means for receiving uploaded audio and video data" means the interface and associated software functionality through which a user provides audio and video data to the system.

[0841] "A speech conversion means for converting speech data into other speech data" refers to the deep learning models and associated processing algorithms used to convert the original speech data into a predetermined speech characteristic or language.

[0842] "Video transformation means for transforming video data into other video data" refers to the deep learning models and associated processing algorithms used to transform specific elements (e.g., facial features) in the original video into predetermined video characteristics or other video elements.

[0843] "Means for synthesizing and outputting converted audio and video data" refers to technology for combining converted audio data and video data to generate a single integrated file or stream and providing it to the user.

[0844] "Means for individually converting and generating personalized educational content" refers to a series of processes for customizing audio and video data for individual learners and creating optimized educational content.

[0845] "Means for delivering educational content" refers to the network infrastructure and related software capabilities required to deliver the generated educational content to students.

[0846] This invention relates to a system for generating and delivering personalized educational content in online education settings, which receives audio and video data uploaded by users, converts the audio and video data into other audio and video data, and synthesizes them to generate optimized educational content.

[0847] The server plays a central role in this system and utilizes the following hardware and software:

[0848] Hardware:

[0849] High-performance processor

[0850] Large capacity memory

[0851] Storage Devices

[0852] Network Interface

[0853] software:

[0854] Operating system (e.g. Linux)

[0855] Deep learning frameworks (e.g., PyTorch)

[0856] Video and audio processing libraries (e.g. FFmpeg, torchaudio)

[0857] Database management system (e.g. MySQL)

[0858] Data processing and calculation:

[0859] 1. Uploading and receiving data:

[0860] Through a web interface, users upload audio and video data for a presentation to a server, which has the means to receive the data and store the audio and video data as separate files.

[0861] 2. Audio data conversion:

[0862] The server uses deep learning models to convert voice data into other voice data, including processes that change voice characteristics and language based on user requests. The deep learning models used include "Speech2TextProcessor" and "Speech2TextForConditionalGeneration," both of which run on PyTorch.

[0863] 3. Video data conversion:

[0864] For video data, a deep learning model is used to transform specific elements (e.g., facial features) based on new facial images provided by the user. FFmpeg and PyTorch are used for video processing.

[0865] 4. Data synthesis and output:

[0866] The converted audio and video data is then recombined and output as new, optimized educational content, which is then distributed to students through educational institutions and online education platforms.

[0867] Examples:

[0868] For example, consider an educational institution that wants to upload audio and video of online lectures and convert them into personalized content in multiple languages. The educational institution's user uploads the original lecture data "lecture_original.mp4" to the system. The following prompt statements are used to specify audio and video conversion:

[0869] Example prompt sentence:

[0870] 1. Voice conversion: "Please translate this audio data into Japanese and resynthesize it in a calm voice."

[0871] 2. Video conversion: "Please replace the speaker's face in the lecture video with another face image."

[0872] The server processes these instructions and ultimately generates an optimized lecture video, "lecture_final.mp4," with the faces modified and translated into Japanese, which it then provides to the educational institution.

[0873] The above format will enable the creation and delivery of educational content suited to each student's learning style and language, which is expected to improve the quality and effectiveness of education.

[0874] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0875] Step 1:

[0876] Users upload audio and video data

[0877] Specific Operation: A user uploads an audio and video data file of a lecture or presentation (e.g., lecture_original.mp4) to the server using the web interface.

[0878] Input: Audio and video data files

[0879] Output: Audio and video data files stored on the server

[0880] Step 2:

[0881] The server receives the uploaded data and performs preprocessing.

[0882] Specific operation: The server splits the received audio and video data and saves them as audio data (e.g., extracted_audio.wav) and video data (e.g., video_no_audio.mp4). At the same time, it uses FFmpeg to perform preprocessing such as noise reduction and resolution conversion.

[0883] Input: Uploaded audio and video data files

[0884] Output: Preprocessed audio and video data files

[0885] Step 3:

[0886] The server converts the audio data

[0887] Specific operation: The server uses a speech conversion method (e.g., Speech2TextProcessor, Speech2TextForConditionalGeneration) to convert the voice data based on the user's prompts, and uses a generative AI model to modify the voice characteristics and language specified, generating a new voice data file (e.g., translated_audio.wav).

[0888] Input: Preprocessed audio data and speech-to-speech prompt

[0889] Output: Converted audio data file

[0890] Step 4:

[0891] The server converts the video data

[0892] Specific operation: The server converts specific elements (e.g., facial features) in the video data using a video conversion method (e.g., Face Swap, etc.) and generates a new video data file (e.g., converted_video.mp4) based on the prompt text.

[0893] Input: Preprocessed video data and video transformation prompt

[0894] Output: Converted video data file

[0895] Step 5:

[0896] The server synthesizes the converted audio and video data.

[0897] Specific operation: The server synthesizes the converted audio data (e.g., translated_audio.wav) and video data (e.g., converted_video.mp4) to generate a new integrated educational content file (e.g., final_presentation.mp4), using FFmpeg to synchronize the audio and video.

[0898] Input: Converted audio and video data files

[0899] Output: Composite educational content file

[0900] Step 6:

[0901] Download or distribute user-generated educational content

[0902] Specific operation: The user downloads the new educational content file (e.g., final_presentation.mp4) generated on the system or distributes it through the educational platform, allowing students to view personalized content.

[0903] Input: Synthesized educational content file

[0904] Output: User download or distribution of content

[0905] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0906] The present invention relates to a system that combines means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, means for synthesizing and outputting the converted audio and video data, and an emotion engine that recognizes user emotions. The system of the present invention utilizes generative AI and the emotion engine to encourage honest expression of opinions and improve communication in the workplace.

[0907] Program processing

[0908] 1. User upload of data

[0909] Users upload presentation audio and video files to the system. This is done using a common web interface. The user selects the files and sends them to the server, where the data is ingested into the system.

[0910] 2. Data preprocessing by the server

[0911] The server receives the uploaded file and separates the audio and video data, saving them as separate files and performing pre-processing such as noise reduction and resolution adjustment if necessary. This step ensures a smooth conversion process.

[0912] 3. Emotional engine for recognizing user emotions

[0913] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's voice intonation, facial expressions, or biometric sensor data to identify the user's emotional state. This information is reflected in the subsequent audio and video conversion process.

[0914] 4. Data transformation by the server

[0915] Based on the results of the emotion engine, the server applies a voice conversion model to convert the original audio data into a target voice. In this case, it uses a deep learning model to convert it into the voice of the project manager's colleague. The converted result is saved as "converted_audio.mp3."

[0916] The server then applies the video conversion model to convert the original video data into a target video, which is then converted into the face of the project manager's colleague using a deep learning model. The converted video data is saved as "converted_video.mp4."

[0917] 5. Data synthesis and output by the server

[0918] The converted audio and video data are then combined by the server and output as a new presentation file, which is saved as "final_presentation.mp4" and provided to the user.

[0919] Specific examples

[0920] For example, imagine a project manager at a company receives an important presentation from an executive, saved in a file called "original_presentation.mp4."

[0921] 1. The project manager (user) uploads "original_presentation.mp4" to the system.

[0922] 2. The server receives this file and separates the audio and video into "original_audio.mp3" and "original_video.mp4" respectively.

[0923] 3. The server uses the emotion engine to recognize the user's emotion, which is reflected in the subsequent conversion process.

[0924] 4. Based on the recognized emotion, the server applies a voice conversion model to convert "original_audio.mp3" into the voice of the project manager's colleague. The converted result is saved as "converted_audio.mp3."

[0925] 5. Next, the server applies the video conversion model to convert "original_video.mp4" into a video showing the colleague's face. The converted video data is saved as "converted_video.mp4."

[0926] 6. Finally, the server combines "converted_audio.mp3" and "converted_video.mp4" to create a new file called "final_presentation.mp4."

[0927] 7. The project manager (user) downloads this "final_presentation.mp4" and watches the presentation with the faces and voices of their colleagues.

[0928] This allows project managers, unlike executives, to listen to presentations in a more relaxed state and more easily express their opinions frankly. This is expected to lead to improved communication and constructive exchange of opinions within the workplace. In this way, the system of the present invention reduces communication barriers caused by differences in position and provides an effective means for improving the work environment.

[0929] The processing flow will be explained below.

[0930] Step 1:

[0931] Users upload presentation audio and video files to the system using a common web interface where the user selects the files and sends them to the server.

[0932] Step 2:

[0933] The server receives the uploaded file and separates the audio and video data. Specifically, the server extracts the audio from the received video file and saves the audio data as "original_audio.mp3" and the video data as "original_video.mp4."

[0934] Step 3:

[0935] The server performs pre-processing such as noise reduction and normalization on audio data, and resolution and frame rate adjustment on video data to ensure the conversion process proceeds efficiently and accurately.

[0936] Step 4:

[0937] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the voice intonation, facial expressions, or biometric sensor data provided when the user accesses the system to identify the user's emotional state. This emotion information is used in the subsequent conversion process.

[0938] Step 5:

[0939] Based on the emotional information recognized by the emotion engine, the server applies a voice conversion model to convert the original voice data into a target voice, for example, the voice of a colleague of the project manager, and the converted voice data is saved as "converted_audio.mp3."

[0940] Step 6:

[0941] Similarly, the server applies the video conversion model based on the emotion information recognized by the emotion engine to convert the original video data into a target video, specifically, the face of the project manager's colleague, and saves the converted video data as "converted_video.mp4."

[0942] Step 7:

[0943] The server combines the converted audio and video data to create a new presentation file. Specifically, it merges "converted_audio.mp3" and "converted_video.mp4" to create a new file called "final_presentation.mp4."

[0944] Step 8:

[0945] Users download "final_presentation.mp4" from the system. They can play this file and watch the presentation being given by their colleagues, with their faces and voices. This process creates an environment where participants can express their opinions frankly without feeling out of place when receiving a presentation from a senior executive.

[0946] Through these steps, the system of the present invention can provide users with an effective means of reducing status differences and improving communication within the workplace.

[0947] Example 2

[0948] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0949] Conventional presentation systems have the problem that it is difficult to exchange opinions frankly due to differences in hierarchical relationships and positions. In particular, constructive exchange of opinions and communication within the workplace is often hindered because the content of the presentation cannot be received in a more relaxed environment. This causes problems such as reduced efficiency and productivity throughout the organization.

[0950] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0951] In this invention, the server includes means for receiving uploaded audio and video data, audio conversion means for converting audio data into other audio data, video conversion means for converting video data into other video data, means for synthesizing and outputting the converted audio and video data, emotion recognition means for recognizing a user's emotion, means for reflecting the emotion information obtained by the emotion recognition means in the audio and video conversion process, and process control means for managing a series of processes from uploading to output. This allows users to receive presentations in different emotional states and facilitates frank exchange of opinions in a relaxed environment.

[0952] "Audio data" means data that records sound and is stored in digital or analog format.

[0953] "Video data" refers to data that records video and is stored in digital or analog format.

[0954] "Means for receiving" refers to devices or software that have the function of acquiring data and incorporating it into the system.

[0955] "Audio conversion means" refers to a device or software that has the function of converting original audio data into another audio data.

[0956] "Video conversion means" refers to a device or software that has the function of converting original video data into other video data.

[0957] "Means for synthesis and output" refers to devices or software that have the function of combining multiple pieces of data (audio, video, etc.) into the final output format.

[0958] "Emotion recognition means" refers to a device or software that has the function of analyzing audio, video, or biometric sensor data to identify a user's emotions.

[0959] "Process control means" refers to devices or software that have the function of managing the entire process in an orderly manner and properly executing each processing step.

[0960] A "machine learning model" refers to an algorithm or technique that learns patterns based on specific data and processes new data according to those patterns.

[0961] The system of the present invention converts uploaded audio and video data into other audio and video data, and finally synthesizes and outputs them. This system has the function of recognizing the user's emotions and reflecting this information in the audio and video conversion process.

[0962] Hardware and Software Configuration

[0963] The system implementation uses the following major hardware and software:

[0964] 1. Server: A central processing unit that receives, preprocesses, converts, synthesizes, and outputs data. Specifically, a server machine with high processing power is used.

[0965] 2. Web Interface: This is an interface via a web browser for users to upload audio and video data. Users access the file upload page in their browser, click the file selection button to select a file, and then press the upload button to send the data to the server.

[0966] 3. Audio analysis software: A library for processing the uploaded audio data, such as FFmpeg or OpenSmile, which separates, preprocesses, and converts the audio data.

[0967] 4. Video analysis software: Libraries for processing the uploaded video data, such as FFmpeg and OpenCV, which separate, preprocess, and convert the video data.

[0968] 5. Emotion Recognition Engine: Software for analyzing a user's emotions from audio and video data. It identifies their emotional state by analyzing their voice intonation and facial expressions.

[0969] 6. Machine learning models: These are models that use deep learning to convert audio and video data. Specifically, models such as WaveNet and FaceSwap are used.

[0970] Example of a system

[0971] For example, imagine a project manager at a company receives an important presentation, which is saved in a file called "original_presentation.mp4."

[0972] 1. The project manager (user) uploads "original_presentation.mp4" to the system. Using the web interface, select the file on the file upload page and submit it.

[0973] 2. The server receives this file and separates the audio and video data using FFmpeg to create audio data "original_audio.mp3" and video data "original_video.mp4".

[0974] 3. The server uses an emotion recognition engine to recognize the user's emotions. It uses OpenSmile and OpenCV to analyze the emotional state of the voice and video, and obtains information such as whether the user is excited or happy.

[0975] 4. The server applies a voice conversion model based on the results of the emotion recognition engine to convert "original_audio.mp3" into the colleague's voice. This conversion uses WaveNet, and the result is saved as "converted_audio.mp3."

[0976] 5. The server applies the video transformation model to convert "original_video.mp4" to the colleague's face. This transformation uses FaceSwap, and the result is saved as "converted_video.mp4."

[0977] 6. The server combines the converted audio and video data and saves it as the final file "final_presentation.mp4." It uses FFmpeg to combine the audio and video to generate a new multimedia file.

[0978] In this way, project managers can see and hear their colleagues' presentations, evaluate and exchange opinions in a relaxed environment.

[0979] Prompt Sentence Examples

[0980] "Please describe a system that uploads presentation audio and video data, uses an emotion recognition engine to convert the audio and video to that of a colleague, and finally generates a composite presentation file."

[0981] This system will enable users to exchange opinions frankly without worrying about hierarchical relationships or differences in position, and is expected to improve communication in the workplace.

[0982] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0983] Step 1:

[0984] The user uploads the presentation audio and video file ("original_presentation.mp4") to the system through a web interface: the user accesses the file upload page in their browser, clicks the file selection button to select the presentation file, and presses the 'Upload' button, which sends the selected file to the server.

[0985] Input: Presentation file "original_presentation.mp4"

[0986] Output: File upload request to the server

[0987] Step 2:

[0988] The server receives the uploaded "original_presentation.mp4" and separates the audio and video data into "original_audio.mp3" and "original_video.mp4" using the FFmpeg library.

[0989] Input: Uploaded presentation file "original_presentation.mp4"

[0990] Output: Separated audio data "original_audio.mp3" and video data "original_video.mp4"

[0991] Step 3:

[0992] The server preprocesses the received audio data "original_audio.mp3" and video data "original_video.mp4" by applying noise reduction to the audio data and adjusting the resolution of the video data.

[0993] Input: Separated audio data "original_audio.mp3", video data "original_video.mp4"

[0994] Output: Preprocessed audio and video data

[0995] Step 4:

[0996] The server uses the pre-processed audio and video data to recognize the user's emotions. It uses an emotion recognition engine to analyze intonation from the audio data and facial expressions from the video data to identify the user's emotional state.

[0997] Input: Preprocessed audio data, preprocessed video data

[0998] Output: User's emotional state data (e.g., excitement, joy)

[0999] Step 5:

[1000] Based on the recognized emotional state, the server applies a voice conversion model to convert the original audio data "original_audio.mp3" into the target audio data "converted_audio.mp3", using a deep learning model such as WaveNet.

[1001] Input: Preprocessed audio data "original_audio.mp3", user emotional state data

[1002] Output: Converted audio data "converted_audio.mp3"

[1003] Step 6:

[1004] Based on the recognized emotional state, the server applies a video conversion model to convert the original video data "original_video.mp4" into the target video data "converted_video.mp4." This conversion uses a deep learning model such as FaceSwap.

[1005] Input: Preprocessed video data "original_video.mp4", user emotional state data

[1006] Output: Converted video data "converted_video.mp4"

[1007] Step 7:

[1008] The server combines the converted audio data (converted_audio.mp3) and video data (converted_video.mp4) and outputs the combined data as a new presentation file (final_presentation.mp4). FFmpeg is used to combine the audio and video to generate a new multimedia file.

[1009] Input: Converted audio data "converted_audio.mp3", converted video data "converted_video.mp4"

[1010] Output: The final composite presentation file "final_presentation.mp4"

[1011] This series of processes allows users to watch presentations delivered in the presence and voice of their colleagues, and to evaluate and exchange opinions in a relaxed environment.

[1012] (Application example 2)

[1013] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1014] Conventional presentation systems have difficulty adapting the audio and video of a presentation to a specific target audience, and are unable to recognize and adapt the content based on user emotions, limiting their effectiveness in attracting audience attention. Furthermore, they lack the means to allow users to relax and express their honest opinions. This invention solves these problems and provides a more effective presentation adaptation and viewing experience.

[1015] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1016] In this invention, the server includes means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, means for synthesizing and outputting the converted audio and video data, emotion recognition means for recognizing a user's emotion, control means for converting the audio and video data based on the emotion information recognized by the emotion recognition means, and means for synthesizing and outputting the converted audio and video data in a viewable format. This allows users to convert uploaded presentations to suit specific targets and make adjustments based on emotions to create more effective and attractive presentations. This also allows viewers to watch the content in a relaxed atmosphere and more easily express their honest opinions.

[1017] "Upload" refers to the operation in which a user sends a local file to the server and imports the data into the system.

[1018] "Presentation audio data" is digital data containing audio information used in a presentation.

[1019] "Presentation video data" is digital data containing video information used in a presentation.

[1020] "Audio conversion means" refers to a function or device for converting uploaded presentation audio data into other audio data.

[1021] "Video conversion means" refers to a function or device for converting uploaded presentation video data into other video data.

[1022] "Emotion recognition means" refers to functions or devices for analyzing and recognizing emotions from the user's voice or video.

[1023] The "control means" refers to a function or device that supervises and controls the conversion of audio and video data based on the emotional information obtained by the emotional recognition means.

[1024] "Synthesis" refers to the process of combining multiple pieces of data into a single format.

[1025] A "deep learning model" is an artificial intelligence algorithm that uses multi-layered neural networks to learn from data and perform specific tasks.

[1026] A "prompt" is a document that contains instructions or questions given to a generative AI model, and is an input document used to obtain specific results.

[1027] The system required to implement this invention includes the following hardware and software. The main hardware required is a smartphone and a server. The software requires a web interface, deep learning models (e.g., TensorFlow, PyTorch), audio and video processing libraries (e.g., WaveNet, GANs), and emotion recognition APIs (e.g., Google Cloud Vision API, Microsoft Azure Emotion API).

[1028] The embodiment of this system mainly operates through the following process. The process begins with the user uploading presentation audio and video data. The server then separates the data and saves them as separate files. The server then uses emotion recognition means to analyze emotions from the user's audio and video. The obtained emotion information is reflected in the audio and video data conversion process.

[1029] As a concrete example, consider the following scenario: A user wants to convert an important internal presentation into a character played by a famous actor. The user uploads the audio and video data of the presentation from their smartphone to the application. This data is sent to the server, where it is separated into audio and video. Next, an emotion recognition API is used to analyze the user's emotions, and a deep learning model is used to convert the audio and video data into that of the target character. This conversion is performed using prompt sentences to provide specific instructions.

[1030] For example, you can use the following prompt:

[1031] "Translate this presentation into the audio and visuals of a famous movie actor, maintaining his signature friendly and calming tone."

[1032] The converted audio and video data is then recombined on the server side and output as a new presentation. Users can then easily view and share the resulting presentation on their smartphones. Through this entire process, users can view the presentation in a more relaxed state and more easily express their honest opinions.

[1033] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1034] Step 1:

[1035] Uploading data

[1036] The user uploads the presentation audio and video data to the application from their smartphone. The user selects the file and initiates the upload using the web interface. The input is the presentation file, and the output is the audio and video data file stored on the server. Specifically, the user selects the target file from the smartphone's file system and taps the send button, which uploads the file to the server.

[1037] Step 2:

[1038] Data Preprocessing

[1039] The server receives the uploaded presentation file and separates the audio and video data. The input is the uploaded presentation file (e.g., .mp4 or .pptx), and the output is separate audio (.mp3) and video (.mp4) files. During this process, preprocessing such as noise removal and resolution adjustment is performed. Specifically, the server analyzes the file and uses the appropriate processing library to separate and save the audio and video.

[1040] Step 3:

[1041] emotion recognition

[1042] The server uses an emotion recognition API to recognize the user's emotions from audio and video data. The input is the separated audio and video files and a prompt from the user, and the output is the recognized emotion data. Specifically, it analyzes the intonation of the voice and the facial features of the video, and obtains an emotion label using the emotion recognition API.

[1043] Step 4:

[1044] Audio data conversion

[1045] The server converts the voice data using a deep learning model based on the emotion recognition results and the prompt. The input is the original voice data, the prompt, and the recognized emotion data, and the output is a converted voice file (converted_audio.mp3) with the target voice characteristics. Specifically, the voice is processed using a voice conversion model such as WaveNet.

[1046] Step 5:

[1047] Video data conversion

[1048] The server converts the video data using a deep learning model based on the emotion recognition results and prompt text. The input is the original video data, the prompt text, and the recognized emotion data, and the output is a converted video file (converted_video.mp4) with the target video characteristics. Specifically, it uses GANs and other tools to process and change faces and specific parts of the video.

[1049] Step 6:

[1050] Data synthesis and output

[1051] The server synthesizes the converted audio and video data and generates a new presentation file (final_presentation.mp4). The input is the converted audio and video files, and the output is the final synthesized presentation file. Specifically, the audio and video are synchronized and synthesized using a multimedia processing library such as FFmpeg.

[1052] Step 7:

[1053] User Viewing and Sharing

[1054] The user downloads the generated presentation to their smartphone, where it can be viewed and shared. The input is the final presentation file, and the output is the user's viewing experience and sharing via social media. Specific actions include using the application to download the video and tapping the play button. The user can also use the share button to share via email or social media.

[1055] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1056] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1057] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1058] [Fourth embodiment]

[1059] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1060] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1061] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1062] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1063] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1064] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1065] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1066] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1067] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1068] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1069] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1070] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1071] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1072] The present invention relates to a system including means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, and means for synthesizing and outputting the converted audio and video data. The system of the present invention utilizes generative AI to encourage candid opinions and improve communication in the workplace.

[1073] Program processing

[1074] 1. User upload of data

[1075] Users upload the presentation's audio and video files to the system, which is done using a common web interface.

[1076] 2. Data preprocessing by the server

[1077] The server receives the uploaded file and separates the audio and video data. The audio and video are saved as separate files, and preprocessing such as noise reduction and resolution conversion is performed as needed.

[1078] 3. Data transformation by the server

[1079] The server uses deep learning models to convert audio and video data: the audio conversion model converts the original audio data into other audio data, and the video conversion model converts the original video data into other video data.

[1080] 4. Data synthesis and output by the server

[1081] The converted audio and video data are then combined by the server and output as a new presentation file, which is then provided to the user.

[1082] Specific examples

[1083] For example, imagine a project manager at a company receives an important presentation from an executive, saved in a file called "original_presentation.mp4."

[1084] 1. The project manager (user) uploads "original_presentation.mp4" to the system.

[1085] 2. The server receives this file and separates the audio and video into "original_audio.mp3" and "original_video.mp4" respectively.

[1086] 3. The server applies the voice conversion model to convert "original_audio.mp3" into the voice of the project manager's colleague, resulting in the audio file "converted_audio.mp3."

[1087] 4. The server then applies the video conversion model to convert "original_video.mp4" into a video showing the colleague's face, resulting in a video file called "converted_video.mp4."

[1088] 5. Finally, the server combines "converted_audio.mp3" and "converted_video.mp4" to create a new file called "final_presentation.mp4."

[1089] 6. The project manager (user) downloads this "final_presentation.mp4" and watches the presentation with the faces and voices of their colleagues.

[1090] This allows project managers, unlike executives, to listen to presentations in a more relaxed state and more easily express their opinions frankly. This is expected to lead to improved communication and constructive exchange of opinions within the workplace. In this way, the system of the present invention reduces communication barriers caused by differences in position and provides an effective means for improving the work environment.

[1091] The processing flow will be explained below.

[1092] Step 1:

[1093] Users upload presentation audio and video files to the system. This process is done using a web interface, where users select files and send them to the server, where they are received and stored.

[1094] Step 2:

[1095] The server receives the uploaded file and separates the audio and video data. The server then reads the received video file, extracts the audio, and saves it as a separate file. Specifically, the audio is saved as "original_audio.mp3" and the video is saved as "original_video.mp4."

[1096] Step 3:

[1097] The server pre-processes the audio data by denoising and normalising it, and also adjusts the resolution and frame rate of the video data as needed, preparing it for a more efficient and accurate conversion process.

[1098] Step 4:

[1099] The server uses a voice conversion model to convert the original audio data into a target voice. In this case, it uses a deep learning model to convert it into the voice of the project manager's colleague. The converted result is saved as "converted_audio.mp3."

[1100] Step 5:

[1101] The server uses a video conversion model to convert the original video data into a target video, in this case the face of the project manager's colleague, using a deep learning model. The converted video data is saved as "converted_video.mp4."

[1102] Step 6:

[1103] The server combines the converted audio and video data. Specifically, it merges "converted_audio.mp3" and "converted_video.mp4" to generate a new presentation file. This new file is saved as "final_presentation.mp4."

[1104] Step 7:

[1105] Users download "final_presentation.mp4" from the system. They can play this file and watch the presentation being given by their colleagues, with their faces and voices. This process creates an environment where participants feel more comfortable listening to a presentation from a senior executive and can express their opinions more frankly.

[1106] Through these steps, the system of the present invention can provide users with an effective means of reducing status differences and improving communication within the workplace.

[1107] Example 1

[1108] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1109] Current presentations rely on the presenter's specific audio and video, making it difficult for viewers to evaluate the content unbiased. Furthermore, when it comes to smooth communication in the workplace, differences in perspective can make it difficult for people to express their opinions frankly. Conventional systems are limited in the technology they use, making it difficult to efficiently convert and synthesize audio and video.

[1110] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1111] In this invention, the server includes means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, means for converting the audio data and video data using a deep learning model, means for performing noise reduction and resolution conversion as preprocessing, and means for synthesizing and outputting the converted audio and video data. This reduces communication barriers due to differences in positions in the workplace and enables constructive exchange of opinions.

[1112] "Upload" is the process by which a user submits data to the system for storage.

[1113] "Presentation audio data" is digital data of the audio used during a presentation.

[1114] "Presentation video data" refers to digital data of the video used during a presentation.

[1115] "Audio conversion means" refers to a method or device for converting original audio data into audio data of a different format or content.

[1116] "Video conversion means" refers to a method or device for converting original video data into video data of a different format or content.

[1117] A "deep learning model" is a machine learning model that uses a multi-layer neural network to convert and analyze data.

[1118] "Preprocessing" refers to the process of performing basic processing such as noise removal and resolution conversion before converting the data.

[1119] "Denoising" is the process of removing unwanted noise from audio data.

[1120] "Resolution conversion" is the process of changing the resolution of video data.

[1121] "Synthesis" is the process of combining multiple pieces of data (audio data or video data) into one piece of data.

[1122] "Output" refers to providing processed data in a form usable by a user.

[1123] The present invention is a system that receives audio and video data of a presentation, converts it, and outputs it. The system processes presentation data uploaded by a user and converts the audio and video data into other formats using a deep learning model, enabling more effective communication.

[1124] First, a user uploads the presentation audio and video files to the system using the web interface. For example, a user may upload a file called "original_presentation.mp4." To do this, the user accesses the specified URL through a web browser and clicks the upload button.

[1125] Next, the server receives the uploaded file and separates the audio and video data. Specifically, it uses a tool such as FFmpeg to save the audio and video as separate files. It then applies a noise reduction filter to the separated audio data and performs resolution conversion on the video data. For example, the audio data is saved as "original_audio.mp3" and the video data as "original_video.mp4."

[1126] The server then uses a deep learning model to convert the audio and video data. It uses an audio conversion model (for example, a model using TensorFlow or PyTorch) to convert the audio data and generate a new audio file called "converted_audio.mp3". Similarly, it uses a video conversion model to generate a new video file called "converted_video.mp4".

[1127] The converted audio and video data are then combined by the server and output as the final presentation file, "final_presentation.mp4." Using a media processing library such as FFmpeg, the audio and video data are combined into a single file. Users can then download this new file and view it in their own environment.

[1128] As a concrete example, consider the case where a project manager at a company receives an important presentation from an executive. This presentation is saved in "original_presentation.mp4." The project manager (user) uploads this file to the system. The server receives this file and separates the audio and video. After noise removal and resolution conversion, it applies audio conversion models and video conversion models to generate new data for each. Finally, it is provided to the user as "final_presentation.mp4."

[1129] Example prompt sentence:

[1130] "Please upload your presentation file. Filename: original_presentation.mp4"

[1131] "Separating audio and video data. Please wait a moment."

[1132] "Converting audio data. Progress: 40%"

[1133] "Converting video data. Progress: 70%"

[1134] "Synthesizing converted data. Progress: 90%"

[1135] "Conversion complete. Download your file. Filename: final_presentation.mp4"

[1136] In this way, the system of the present invention uses deep learning technology and various media processing tools to effectively convert and synthesize presentation data uploaded by users, thereby improving communication within the workplace.

[1137] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1138] Step 1: User uploads data

[1139] A user uploads a presentation's audio and video files to the system using a web interface. The input is a file called "original_presentation.mp4," which is sent to the server. Specifically, the user opens a web browser, accesses the specified URL, clicks the "Choose File" button, selects the desired file, and presses the "Upload" button.

[1140] input:

[1141] Upload file: original_presentation.mp4

[1142] output:

[1143] The file is sent to the server

[1144] Step 2: Server receives and separates data

[1145] The server receives the uploaded "original_presentation.mp4" and uses FFmpeg to separate the audio and video data. Specifically, it saves the received file in a specific directory and executes the FFmpeg command to generate the audio data "original_audio.mp3" and the video data "original_video.mp4."

[1146] input:

[1147] Upload file: original_presentation.mp4

[1148] output:

[1149] Separated audio data: original_audio.mp3

[1150] Separated video data: original_video.mp4

[1151] Step 3: Preprocessing the data on the server

[1152] The server performs preprocessing on the separated audio and video data, applying an audio filter to remove noise and converting the resolution. Specifically, it uses the FFmpeg command to remove noise from the audio data (e.g., ffmpeg -i original_audio.mp3 -af "highpass=f=200, lowpass=f=3000" cleaned_audio.mp3) and converts the resolution of the video data (e.g., ffmpeg -i original_video.mp4 -vf scale=1280:720 resized_video.mp4).

[1153] input:

[1154] Audio data: original_audio.mp3

[1155] Video data: original_video.mp4

[1156] output:

[1157] Noise-removed audio data: cleaned_audio.mp3

[1158] Resolution converted video data: resized_video.mp4

[1159] Step 4: Data transformation by the server

[1160] The server converts audio and video data using a deep learning model. Specifically, it inputs audio data into an audio conversion model using TensorFlow or PyTorch, generating a converted audio file called "converted_audio.mp3." Similarly, it inputs video data into a video conversion model and generates a converted video file called "converted_video.mp4" (e.g., python audio_conversion.py --input cleaned_audio.mp3 --output converted_audio.mp3 and python video_conversion.py --input resized_video.mp4 --output converted_video.mp4).

[1161] input:

[1162] Preprocessed audio data: cleaned_audio.mp3

[1163] Preprocessed video data: resized_video.mp4

[1164] output:

[1165] Converted audio data: converted_audio.mp3

[1166] Converted video data: converted_video.mp4

[1167] Step 5: Data synthesis by the server

[1168] The server combines the converted audio and video data to generate the final presentation file, "final_presentation.mp4." Specifically, it uses FFmpeg to execute a command to combine the audio and video into a single file (e.g., ffmpeg -i converted_video.mp4 -i converted_audio.mp3 -c:v copy -c:a aac final_presentation.mp4).

[1169] input:

[1170] Converted audio data: converted_audio.mp3

[1171] Converted video data: converted_video.mp4

[1172] output:

[1173] Final presentation file: final_presentation.mp4

[1174] Step 6: User downloads final file

[1175] The user downloads the generated final presentation file, “final_presentation.mp4,” via the web interface by clicking the download link provided by the system.

[1176] input:

[1177] Final presentation file: final_presentation.mp4

[1178] output:

[1179] Downloaded file: final_presentation.mp4

[1180] (Application example 1)

[1181] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1182] With the spread of online education, there is a demand for improving the quality and effectiveness of educational content. However, traditional educational content has difficulty adapting to different learning styles and languages, which can result in reduced learning effectiveness. In particular, providing content in multiple languages ​​and at an individual learning pace is a challenge.

[1183] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1184] In this invention, the server includes means for receiving uploaded audio and video data, audio conversion means for converting audio data into other audio data, and video conversion means for converting video data into other video data, thereby enabling the audio and video data to be converted individually and personalized educational content to be generated and distributed.

[1185] "Means for receiving uploaded audio and video data" means the interface and associated software functionality through which a user provides audio and video data to the system.

[1186] "A speech conversion means for converting speech data into other speech data" refers to the deep learning models and associated processing algorithms used to convert the original speech data into a predetermined speech characteristic or language.

[1187] "Video transformation means for transforming video data into other video data" refers to the deep learning models and associated processing algorithms used to transform specific elements (e.g., facial features) in the original video into predetermined video characteristics or other video elements.

[1188] "Means for synthesizing and outputting converted audio and video data" refers to technology for combining converted audio data and video data to generate a single integrated file or stream and providing it to the user.

[1189] "Means for individually converting and generating personalized educational content" refers to a series of processes for customizing audio and video data for individual learners and creating optimized educational content.

[1190] "Means for delivering educational content" refers to the network infrastructure and related software capabilities required to deliver the generated educational content to students.

[1191] This invention relates to a system for generating and delivering personalized educational content in online education settings, which receives audio and video data uploaded by users, converts the audio and video data into other audio and video data, and synthesizes them to generate optimized educational content.

[1192] The server plays a central role in this system and utilizes the following hardware and software:

[1193] Hardware:

[1194] High-performance processor

[1195] Large capacity memory

[1196] Storage Devices

[1197] Network Interface

[1198] software:

[1199] Operating system (e.g. Linux)

[1200] Deep learning frameworks (e.g., PyTorch)

[1201] Video and audio processing libraries (e.g. FFmpeg, torchaudio)

[1202] Database management system (e.g. MySQL)

[1203] Data processing and calculation:

[1204] 1. Uploading and receiving data:

[1205] Through a web interface, users upload audio and video data for a presentation to a server, which has the means to receive the data and store the audio and video data as separate files.

[1206] 2. Audio data conversion:

[1207] The server uses deep learning models to convert voice data into other voice data, including processes that change voice characteristics and language based on user requests. The deep learning models used include "Speech2TextProcessor" and "Speech2TextForConditionalGeneration," both of which run on PyTorch.

[1208] 3. Video data conversion:

[1209] For video data, a deep learning model is used to transform specific elements (e.g., facial features) based on new facial images provided by the user. FFmpeg and PyTorch are used for video processing.

[1210] 4. Data synthesis and output:

[1211] The converted audio and video data is then recombined and output as new, optimized educational content, which is then distributed to students through educational institutions and online education platforms.

[1212] Examples:

[1213] For example, consider an educational institution that wants to upload audio and video of online lectures and convert them into personalized content in multiple languages. The educational institution's user uploads the original lecture data "lecture_original.mp4" to the system. The following prompt statements are used to specify audio and video conversion:

[1214] Example prompt sentence:

[1215] 1. Voice conversion: "Please translate this audio data into Japanese and resynthesize it in a calm voice."

[1216] 2. Video conversion: "Please replace the speaker's face in the lecture video with another face image."

[1217] The server processes these instructions and ultimately generates an optimized lecture video, "lecture_final.mp4," with the faces modified and translated into Japanese, which it then provides to the educational institution.

[1218] The above format will enable the creation and delivery of educational content suited to each student's learning style and language, which is expected to improve the quality and effectiveness of education.

[1219] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1220] Step 1:

[1221] Users upload audio and video data

[1222] Specific Operation: A user uploads an audio and video data file of a lecture or presentation (e.g., lecture_original.mp4) to the server using the web interface.

[1223] Input: Audio and video data files

[1224] Output: Audio and video data files stored on the server

[1225] Step 2:

[1226] The server receives the uploaded data and performs preprocessing.

[1227] Specific operation: The server splits the received audio and video data and saves them as audio data (e.g., extracted_audio.wav) and video data (e.g., video_no_audio.mp4). At the same time, it uses FFmpeg to perform preprocessing such as noise reduction and resolution conversion.

[1228] Input: Uploaded audio and video data files

[1229] Output: Preprocessed audio and video data files

[1230] Step 3:

[1231] The server converts the audio data

[1232] Specific operation: The server uses a speech conversion method (e.g., Speech2TextProcessor, Speech2TextForConditionalGeneration) to convert the voice data based on the user's prompts, and uses a generative AI model to modify the voice characteristics and language specified, generating a new voice data file (e.g., translated_audio.wav).

[1233] Input: Preprocessed audio data and speech-to-speech prompt

[1234] Output: Converted audio data file

[1235] Step 4:

[1236] The server converts the video data

[1237] Specific operation: The server converts specific elements (e.g., facial features) in the video data using a video conversion method (e.g., Face Swap, etc.) and generates a new video data file (e.g., converted_video.mp4) based on the prompt text.

[1238] Input: Preprocessed video data and video transformation prompt

[1239] Output: Converted video data file

[1240] Step 5:

[1241] The server synthesizes the converted audio and video data.

[1242] Specific operation: The server synthesizes the converted audio data (e.g., translated_audio.wav) and video data (e.g., converted_video.mp4) to generate a new integrated educational content file (e.g., final_presentation.mp4), using FFmpeg to synchronize the audio and video.

[1243] Input: Converted audio and video data files

[1244] Output: Composite educational content file

[1245] Step 6:

[1246] Download or distribute user-generated educational content

[1247] Specific operation: The user downloads the new educational content file (e.g., final_presentation.mp4) generated on the system or distributes it through the educational platform, allowing students to view personalized content.

[1248] Input: Synthesized educational content file

[1249] Output: User download or distribution of content

[1250] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1251] The present invention relates to a system that combines means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, means for synthesizing and outputting the converted audio and video data, and an emotion engine that recognizes user emotions. The system of the present invention utilizes generative AI and the emotion engine to encourage honest expression of opinions and improve communication in the workplace.

[1252] Program processing

[1253] 1. User upload of data

[1254] Users upload presentation audio and video files to the system. This is done using a common web interface. The user selects the files and sends them to the server, where the data is ingested into the system.

[1255] 2. Data preprocessing by the server

[1256] The server receives the uploaded file and separates the audio and video data, saving them as separate files and performing pre-processing such as noise reduction and resolution adjustment if necessary. This step ensures a smooth conversion process.

[1257] 3. Emotion engine for recognizing user emotions

[1258] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's voice intonation, facial expressions, or biometric sensor data to identify the user's emotional state. This information is reflected in the subsequent audio and video conversion process.

[1259] 4. Data transformation by the server

[1260] Based on the results of the emotion engine, the server applies a voice conversion model to convert the original audio data into a target voice. In this case, it uses a deep learning model to convert it into the voice of the project manager's colleague. The converted result is saved as "converted_audio.mp3."

[1261] The server then applies the video conversion model to convert the original video data into a target video, which is then converted into the face of the project manager's colleague using a deep learning model. The converted video data is saved as "converted_video.mp4."

[1262] 5. Data synthesis and output by the server

[1263] The converted audio and video data are then combined by the server and output as a new presentation file, which is saved as "final_presentation.mp4" and provided to the user.

[1264] Specific examples

[1265] For example, imagine a project manager at a company receives an important presentation from an executive, saved in a file called "original_presentation.mp4."

[1266] 1. The project manager (user) uploads "original_presentation.mp4" to the system.

[1267] 2. The server receives this file and separates the audio and video into "original_audio.mp3" and "original_video.mp4" respectively.

[1268] 3. The server uses the emotion engine to recognize the user's emotion, which is reflected in the subsequent conversion process.

[1269] 4. Based on the recognized emotion, the server applies a voice conversion model to convert "original_audio.mp3" into the voice of the project manager's colleague. The converted result is saved as "converted_audio.mp3."

[1270] 5. Next, the server applies the video conversion model to convert "original_video.mp4" into a video showing the colleague's face. The converted video data is saved as "converted_video.mp4."

[1271] 6. Finally, the server combines "converted_audio.mp3" and "converted_video.mp4" to create a new file called "final_presentation.mp4."

[1272] 7. The project manager (user) downloads this "final_presentation.mp4" and watches the presentation with the faces and voices of their colleagues.

[1273] This allows project managers, unlike executives, to listen to presentations in a more relaxed state and more easily express their opinions frankly. This is expected to lead to improved communication and constructive exchange of opinions within the workplace. In this way, the system of the present invention reduces communication barriers caused by differences in position and provides an effective means for improving the work environment.

[1274] The processing flow will be explained below.

[1275] Step 1:

[1276] Users upload presentation audio and video files to the system using a common web interface where the user selects the files and sends them to the server.

[1277] Step 2:

[1278] The server receives the uploaded file and separates the audio and video data. Specifically, the server extracts the audio from the received video file and saves the audio data as "original_audio.mp3" and the video data as "original_video.mp4."

[1279] Step 3:

[1280] The server performs pre-processing such as noise reduction and normalization on audio data, and resolution and frame rate adjustment on video data to ensure the conversion process proceeds efficiently and accurately.

[1281] Step 4:

[1282] The server uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the voice intonation, facial expressions, or biometric sensor data provided when the user accesses the system to identify the user's emotional state. This emotion information is used in the subsequent conversion process.

[1283] Step 5:

[1284] Based on the emotional information recognized by the emotion engine, the server applies a voice conversion model to convert the original voice data into a target voice, for example, the voice of a colleague of the project manager, and the converted voice data is saved as "converted_audio.mp3."

[1285] Step 6:

[1286] Similarly, the server applies the video conversion model based on the emotion information recognized by the emotion engine to convert the original video data into a target video, specifically, the face of the project manager's colleague, and saves the converted video data as "converted_video.mp4."

[1287] Step 7:

[1288] The server combines the converted audio and video data to create a new presentation file. Specifically, it merges "converted_audio.mp3" and "converted_video.mp4" to create a new file called "final_presentation.mp4."

[1289] Step 8:

[1290] Users download "final_presentation.mp4" from the system. They can play this file and watch the presentation being given by their colleagues, with their faces and voices. This process creates an environment where participants can express their opinions frankly without feeling out of place when receiving a presentation from a senior executive.

[1291] Through these steps, the system of the present invention can provide users with an effective means of reducing status differences and improving communication within the workplace.

[1292] Example 2

[1293] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1294] Conventional presentation systems have the problem that it is difficult to exchange opinions frankly due to differences in hierarchical relationships and positions. In particular, constructive exchange of opinions and communication within the workplace is often hindered because the content of the presentation cannot be received in a more relaxed environment. This causes problems such as reduced efficiency and productivity throughout the organization.

[1295] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1296] In this invention, the server includes means for receiving uploaded audio and video data, audio conversion means for converting audio data into other audio data, video conversion means for converting video data into other video data, means for synthesizing and outputting the converted audio and video data, emotion recognition means for recognizing a user's emotion, means for reflecting the emotion information obtained by the emotion recognition means in the audio and video conversion process, and process control means for managing a series of processes from uploading to output. This allows users to receive presentations in different emotional states and facilitates frank exchange of opinions in a relaxed environment.

[1297] "Audio data" means data that records sound and is stored in digital or analog format.

[1298] "Video data" refers to data that records video and is stored in digital or analog format.

[1299] "Means for receiving" refers to devices or software that have the function of acquiring data and incorporating it into the system.

[1300] "Audio conversion means" refers to a device or software that has the function of converting original audio data into another audio data.

[1301] "Video conversion means" refers to a device or software that has the function of converting original video data into other video data.

[1302] "Means for synthesis and output" refers to devices or software that have the function of combining multiple pieces of data (audio, video, etc.) into the final output format.

[1303] "Emotion recognition means" refers to a device or software that has the function of analyzing audio, video, or biometric sensor data to identify a user's emotions.

[1304] "Process control means" refers to devices or software that have the function of managing the entire process in an orderly manner and properly executing each processing step.

[1305] A "machine learning model" refers to an algorithm or technique that learns patterns based on specific data and processes new data according to those patterns.

[1306] The system of the present invention converts uploaded audio and video data into other audio and video data, and finally synthesizes and outputs them. This system has the function of recognizing the user's emotions and reflecting this information in the audio and video conversion process.

[1307] Hardware and Software Configuration

[1308] The system implementation uses the following major hardware and software:

[1309] 1. Server: A central processing unit that receives, preprocesses, converts, synthesizes, and outputs data. Specifically, a server machine with high processing power is used.

[1310] 2. Web Interface: This is an interface via a web browser for users to upload audio and video data. Users access the file upload page in their browser, click the file selection button to select a file, and then press the upload button to send the data to the server.

[1311] 3. Audio analysis software: A library for processing the uploaded audio data, such as FFmpeg or OpenSmile, which separates, preprocesses, and converts the audio data.

[1312] 4. Video analysis software: Libraries for processing the uploaded video data, such as FFmpeg and OpenCV, which separate, preprocess, and convert the video data.

[1313] 5. Emotion Recognition Engine: Software for analyzing a user's emotions from audio and video data. It identifies their emotional state by analyzing their voice intonation and facial expressions.

[1314] 6. Machine learning models: These are models that use deep learning to convert audio and video data. Specifically, models such as WaveNet and FaceSwap are used.

[1315] Example of a system

[1316] For example, imagine a project manager at a company receives an important presentation, which is saved in a file called "original_presentation.mp4."

[1317] 1. The project manager (user) uploads "original_presentation.mp4" to the system. Using the web interface, select the file on the file upload page and submit it.

[1318] 2. The server receives this file and separates the audio and video data using FFmpeg to create audio data "original_audio.mp3" and video data "original_video.mp4".

[1319] 3. The server uses an emotion recognition engine to recognize the user's emotions. It uses OpenSmile and OpenCV to analyze the emotional state of the voice and video, and obtains information such as whether the user is excited or happy.

[1320] 4. The server applies a voice conversion model based on the results of the emotion recognition engine to convert "original_audio.mp3" into the colleague's voice. This conversion uses WaveNet, and the result is saved as "converted_audio.mp3."

[1321] 5. The server applies the video transformation model to convert "original_video.mp4" to the colleague's face. This transformation uses FaceSwap, and the result is saved as "converted_video.mp4."

[1322] 6. The server combines the converted audio and video data and saves it as the final file "final_presentation.mp4." It uses FFmpeg to combine the audio and video to generate a new multimedia file.

[1323] In this way, project managers can see and hear their colleagues' presentations, evaluate and exchange opinions in a relaxed environment.

[1324] Prompt Sentence Examples

[1325] "Please describe a system that uploads presentation audio and video data, uses an emotion recognition engine to convert the audio and video to that of a colleague, and finally generates a composite presentation file."

[1326] This system will enable users to exchange opinions frankly without worrying about hierarchical relationships or differences in position, and is expected to improve communication in the workplace.

[1327] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1328] Step 1:

[1329] The user uploads the presentation audio and video file ("original_presentation.mp4") to the system through a web interface: the user accesses the file upload page in their browser, clicks the file selection button to select the presentation file, and presses the 'Upload' button, which sends the selected file to the server.

[1330] Input: Presentation file "original_presentation.mp4"

[1331] Output: File upload request to the server

[1332] Step 2:

[1333] The server receives the uploaded "original_presentation.mp4" and separates the audio and video data into "original_audio.mp3" and "original_video.mp4" using the FFmpeg library.

[1334] Input: Uploaded presentation file "original_presentation.mp4"

[1335] Output: Separated audio data "original_audio.mp3" and video data "original_video.mp4"

[1336] Step 3:

[1337] The server preprocesses the received audio data "original_audio.mp3" and video data "original_video.mp4" by applying noise reduction to the audio data and adjusting the resolution of the video data.

[1338] Input: Separated audio data "original_audio.mp3", video data "original_video.mp4"

[1339] Output: Preprocessed audio and video data

[1340] Step 4:

[1341] The server uses the pre-processed audio and video data to recognize the user's emotions. It uses an emotion recognition engine to analyze intonation from the audio data and facial expressions from the video data to identify the user's emotional state.

[1342] Input: Preprocessed audio data, preprocessed video data

[1343] Output: User's emotional state data (e.g., excitement, joy)

[1344] Step 5:

[1345] Based on the recognized emotional state, the server applies a voice conversion model to convert the original audio data "original_audio.mp3" into the target audio data "converted_audio.mp3", using a deep learning model such as WaveNet.

[1346] Input: Preprocessed audio data "original_audio.mp3", user emotional state data

[1347] Output: Converted audio data "converted_audio.mp3"

[1348] Step 6:

[1349] Based on the recognized emotional state, the server applies a video conversion model to convert the original video data "original_video.mp4" into the target video data "converted_video.mp4." This conversion uses a deep learning model such as FaceSwap.

[1350] Input: Preprocessed video data "original_video.mp4", user emotional state data

[1351] Output: Converted video data "converted_video.mp4"

[1352] Step 7:

[1353] The server combines the converted audio data (converted_audio.mp3) and video data (converted_video.mp4) and outputs the combined data as a new presentation file (final_presentation.mp4). FFmpeg is used to combine the audio and video to generate a new multimedia file.

[1354] Input: Converted audio data "converted_audio.mp3", converted video data "converted_video.mp4"

[1355] Output: The final composite presentation file "final_presentation.mp4"

[1356] This series of processes allows users to watch presentations delivered in the presence and voice of their colleagues, and to evaluate and exchange opinions in a relaxed environment.

[1357] (Application example 2)

[1358] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1359] Conventional presentation systems have difficulty adapting the audio and video of a presentation to a specific target audience, and are unable to recognize and adapt the content based on user emotions, limiting their effectiveness in attracting audience attention. Furthermore, they lack the means to allow users to relax and express their honest opinions. This invention solves these problems and provides a more effective presentation adaptation and viewing experience.

[1360] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1361] In this invention, the server includes means for receiving uploaded presentation audio and video data, audio conversion means for converting the presentation audio data into other audio data, video conversion means for converting the presentation video data into other video data, means for synthesizing and outputting the converted audio and video data, emotion recognition means for recognizing a user's emotion, control means for converting the audio and video data based on the emotion information recognized by the emotion recognition means, and means for synthesizing and outputting the converted audio and video data in a viewable format. This allows users to convert uploaded presentations to suit specific targets and make adjustments based on emotions to create more effective and attractive presentations. This also allows viewers to watch the content in a relaxed atmosphere and more easily express their honest opinions.

[1362] "Upload" refers to the operation in which a user sends a local file to the server and imports the data into the system.

[1363] "Presentation audio data" is digital data containing audio information used in a presentation.

[1364] "Presentation video data" is digital data containing video information used in a presentation.

[1365] "Audio conversion means" refers to a function or device for converting uploaded presentation audio data into other audio data.

[1366] "Video conversion means" refers to a function or device for converting uploaded presentation video data into other video data.

[1367] "Emotion recognition means" refers to functions or devices for analyzing and recognizing emotions from the user's voice or video.

[1368] The "control means" refers to a function or device that supervises and controls the conversion of audio and video data based on the emotional information obtained by the emotional recognition means.

[1369] "Synthesis" refers to the process of combining multiple pieces of data into a single format.

[1370] A "deep learning model" is an artificial intelligence algorithm that uses multi-layered neural networks to learn from data and perform specific tasks.

[1371] A "prompt" is a document that contains instructions or questions given to a generative AI model, and is an input document used to obtain specific results.

[1372] The system required to implement this invention includes the following hardware and software. The main hardware required is a smartphone and a server. The software requires a web interface, deep learning models (e.g., TensorFlow, PyTorch), audio and video processing libraries (e.g., WaveNet, GANs), and emotion recognition APIs (e.g., Google Cloud Vision API, Microsoft Azure Emotion API).

[1373] The embodiment of this system mainly operates through the following process. The process begins with the user uploading presentation audio and video data. The server then separates the data and saves them as separate files. The server then uses emotion recognition means to analyze emotions from the user's audio and video. The obtained emotion information is reflected in the audio and video data conversion process.

[1374] As a concrete example, consider the following scenario: A user wants to convert an important internal presentation into a character played by a famous actor. The user uploads the audio and video data of the presentation from their smartphone to the application. This data is sent to the server, where it is separated into audio and video. Next, an emotion recognition API is used to analyze the user's emotions, and a deep learning model is used to convert the audio and video data into that of the target character. This conversion is performed using prompt sentences to provide specific instructions.

[1375] For example, you can use the following prompt:

[1376] "Translate this presentation into the audio and visuals of a famous movie actor, maintaining his signature friendly and calming tone."

[1377] The converted audio and video data is then recombined on the server side and output as a new presentation. Users can then easily view and share the resulting presentation on their smartphones. Through this entire process, users can view the presentation in a more relaxed state and more easily express their honest opinions.

[1378] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1379] Step 1:

[1380] Uploading data

[1381] The user uploads the presentation audio and video data to the application from their smartphone. The user selects the file and initiates the upload using the web interface. The input is the presentation file, and the output is the audio and video data file stored on the server. Specifically, the user selects the target file from the smartphone's file system and taps the send button, which uploads the file to the server.

[1382] Step 2:

[1383] Data Preprocessing

[1384] The server receives the uploaded presentation file and separates the audio and video data. The input is the uploaded presentation file (e.g., .mp4 or .pptx), and the output is separate audio (.mp3) and video (.mp4) files. During this process, preprocessing such as noise removal and resolution adjustment is performed. Specifically, the server analyzes the file and uses the appropriate processing library to separate and save the audio and video.

[1385] Step 3:

[1386] emotion recognition

[1387] The server uses an emotion recognition API to recognize the user's emotions from audio and video data. The input is the separated audio and video files and a prompt from the user, and the output is the recognized emotion data. Specifically, it analyzes the intonation of the voice and the facial features of the video, and obtains an emotion label using the emotion recognition API.

[1388] Step 4:

[1389] Audio data conversion

[1390] The server converts the voice data using a deep learning model based on the emotion recognition results and the prompt. The input is the original voice data, the prompt, and the recognized emotion data, and the output is a converted voice file (converted_audio.mp3) with the target voice characteristics. Specifically, the voice is processed using a voice conversion model such as WaveNet.

[1391] Step 5:

[1392] Video data conversion

[1393] The server converts the video data using a deep learning model based on the emotion recognition results and prompt text. The input is the original video data, the prompt text, and the recognized emotion data, and the output is a converted video file (converted_video.mp4) with the target video characteristics. Specifically, it uses GANs and other tools to process and change faces and specific parts of the video.

[1394] Step 6:

[1395] Data synthesis and output

[1396] The server synthesizes the converted audio and video data and generates a new presentation file (final_presentation.mp4). The input is the converted audio and video files, and the output is the final synthesized presentation file. Specifically, the audio and video are synchronized and synthesized using a multimedia processing library such as FFmpeg.

[1397] Step 7:

[1398] User Viewing and Sharing

[1399] The user downloads the generated presentation to their smartphone, where it can be viewed and shared. The input is the final presentation file, and the output is the user's viewing experience and sharing via social media. Specific actions include using the application to download the video and tapping the play button. The user can also use the share button to share via email or social media.

[1400] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1401] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1402] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1403] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1404] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1405] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1406] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1407] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1408] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1409] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1410] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1411] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1412] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1413] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1414] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1415] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1416] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1417] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1418] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1419] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1420] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1421] The following is further disclosed regarding the above embodiment.

[1422] (Claim 1)

[1423] means for receiving uploaded presentation audio and video data;

[1424] an audio conversion means for converting the presentation audio data into other audio data;

[1425] a video conversion means for converting the presentation video data into other video data;

[1426] a system including means for synthesizing and outputting the converted audio and video data.

[1427] (Claim 2)

[1428] 10. The system of claim 1, wherein the voice conversion means converts voice characteristics using a deep learning model.

[1429] (Claim 3)

[1430] 10. The system of claim 1, wherein the video transformation means transforms facial features in the video using a deep learning model.

[1431]

[1432] "Example 1"

[1433] (Claim 1)

[1434] means for receiving uploaded presentation audio and video data;

[1435] an audio conversion means for converting the presentation audio data into other audio data;

[1436] a video conversion means for converting the presentation video data into other video data;

[1437] means for synthesizing and outputting the converted audio and video data;

[1438] a means for converting audio and video data utilizing a deep learning model;

[1439] The system includes means for performing noise removal and resolution conversion as the preprocessing.

[1440] (Claim 2)

[1441] The system of claim 1, wherein the voice conversion means converts voice characteristics and converts voice data using deep learning technology.

[1442] (Claim 3)

[1443] 2. The system of claim 1, wherein the image conversion means converts facial features in the image and converts the image data using deep learning technology.

[1444] "Application Example 1"

[1445] (Claim 1)

[1446] means for receiving the uploaded audio and video data;

[1447] a voice conversion means for converting the voice data into other voice data;

[1448] a video conversion means for converting the video data into other video data;

[1449] means for synthesizing and outputting the converted audio and video data;

[1450] means for individually converting the audio data and the video data to generate personalized educational content;

[1451] A system including means for delivering said educational content.

[1452] (Claim 2)

[1453] 10. The system of claim 1, wherein the voice conversion means converts voice characteristics using a deep learning model.

[1454] (Claim 3)

[1455] 10. The system of claim 1, wherein the video transformation means transforms facial features in the video using a deep learning model.

[1456] "Example 2: Combining Emotion Engines"

[1457] (Claim 1)

[1458] means for receiving the uploaded audio and video data;

[1459] a voice conversion means for converting the voice data into other voice data;

[1460] a video conversion means for converting the video data into other video data;

[1461] means for synthesizing and outputting the converted audio and video data;

[1462] emotion recognition means for recognizing an emotion of a user;

[1463] a means for reflecting the emotion information obtained by the emotion recognition means in a voice and video conversion process;

[1464] A system including a process control means for managing a series of processes from upload to output.

[1465] (Claim 2)

[1466] 10. The system of claim 1, wherein the voice conversion means converts voice characteristics using a machine learning model.

[1467] (Claim 3)

[1468] 10. The system of claim 1, wherein the video transformation means transforms facial features in the video using a machine learning model.

[1469] "Application example 2 when combining emotion engines"

[1470] (Claim 1)

[1471] means for receiving uploaded presentation audio and video data;

[1472] an audio conversion means for converting the presentation audio data into other audio data;

[1473] a video conversion means for converting the presentation video data into other video data;

[1474] means for synthesizing and outputting the converted audio and video data;

[1475] emotion recognition means for recognizing an emotion of a user;

[1476] a control means for converting audio and video data based on the emotion information recognized by the emotion recognition means;

[1477] a means for synthesizing the converted audio and video data and outputting the synthesized audio and video data in a viewable format;

[1478] (Claim 2)

[1479] 2. The system of claim 1, wherein the voice conversion means converts voice characteristics using a deep learning model to convert into a target voice based on a prompt sentence.

[1480] (Claim 3)

[1481] 2. The system of claim 1, wherein the video conversion means converts facial features in the video using a deep learning model and converts them into a target video based on a prompt sentence. [Explanation of symbols]

[1482] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving uploaded presentation audio and video data; an audio conversion means for converting the presentation audio data into other audio data; a video conversion means for converting the presentation video data into other video data; a system including means for synthesizing and outputting the converted audio and video data.

2. The system of claim 1 , wherein the voice conversion means converts voice characteristics using a deep learning model.

3. The system of claim 1 , wherein the video transformation means transforms facial features in the video using a deep learning model.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A