System
The system uses generative AI to enhance and reformat video and audio from old videotapes, addressing quality issues and enabling playback on modern devices.
Patent Information
- Application Number
- JP2024118206
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-02-04
AI Technical Summary
Old videotapes, particularly VHS and miniDV, suffer from deteriorating image and audio quality, making them difficult to play on modern devices and lacking effective digitization methods for high-quality digital conversion.
A system using generative AI models to enhance image and audio quality, followed by reformatting the data into user-specified formats for playback on modern devices.
Converts old videotapes into high-quality digital data that can be easily played and shared on modern devices, improving the viewing experience.
Smart Images

Figure 2026017424000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Video and audio stored on old videotapes tend to deteriorate over time, resulting in a poor viewing experience. Video recorded on analog media such as VHS and miniDV is particularly prone to low image quality and noise and blurring. Furthermore, these videotapes have not yet been digitized, making them difficult to play on modern DVDs or smartphones. For this reason, there is a demand for converting these videos to high-quality digital data that can be easily played and shared. [Means for solving the problem]
[0005] To solve this problem, the present invention provides the following means. First, a means for receiving digitized video data from old videotapes is provided. Next, a means for improving the image quality of the video portion of the digitized video data using a generative AI model is provided. Furthermore, a means for improving the sound quality of the audio portion of the same digital data using a generative AI model is provided. Finally, a means for reformatting the video data with improved image quality and sound quality into a user-specified format is provided. A means for providing this high-quality digitized video data to a user is also provided. This converts the video from old videotapes into a state that can be easily played on modern devices, significantly improving the viewing experience.
[0006] "Videotape" is an analog recording medium that uses magnetic tape to record and play back video and audio.
[0007] "Digital video data" refers to video and audio recorded in analog format that has been converted into electronic data format.
[0008] A "generative AI model" is a computer program designed based on deep learning technology and containing algorithms that improve the quality of video and audio data (improving image quality and sound quality).
[0009] "High-definition" refers to the process of improving the resolution of an image, removing noise and blur, and performing color correction.
[0010] "High-quality sound" refers to the process of removing noise from audio data and performing processes such as pitch adjustment and clearing to improve sound quality.
[0011] "Reformatting" is the process of converting processed video and audio data into another specified file format.
[0012] "Means for providing to users" refers to the technical means for providing high-quality video data and high-quality audio in a format that is accessible to users. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] The present invention is a system for converting video data digitized from old videotapes into high-quality image and sound using generation AI, and dubbing it into a format that can be played on DVDs, smartphones, etc. The main modes for implementing the present invention are described in detail below.
[0035] Uploading video data
[0036] Users convert old videotapes (e.g., VHS or miniDV) into digital files using a dedicated digitizing device and save them on their device. Next, they open a specific website or application on their device to upload the saved digital data to a server. The web page provides an interface for selecting and uploading files.
[0037] Example: A user uses a home video capture device to digitize their VHS tapes and save them on their computer. Then, they access a dedicated website, select the saved video files, and upload them.
[0038] High-quality image and sound processing
[0039] The server receives and temporarily stores the uploaded video file. It then processes the received data using a generative AI model. The generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs noise reduction and clearing on the audio track to improve sound quality.
[0040] Example: A generative AI model upscales an original 480p home video to 1080p and produces a noise-reduced video. The audio track is processed to remove background noise and improve audio clarity.
[0041] Reformatting and providing data
[0042] Once the processing is complete, the server converts the high-quality video file to the format of the user's choice (e.g. DVD, MP4, etc.). The converted data is then re-saved and a download link is generated. This link is then communicated to the user for easy access.
[0043] Users can download the converted video data using the provided link, burn it to a disc for playback on a home DVD player, play it on a smartphone, or share it with family and friends using a communication application or cloud service for data sharing.
[0044] For example, a user downloads an ISO file provided by a server and burns it to a DVD disc using home DVD burning software. Then, the user plays the DVD on a home DVD player and watches it with his or her family. Alternatively, the user uploads the converted MP4 file to cloud storage and shares it with family and friends using a specific communication application.
[0045] Thus, the present invention is a system that uses generative AI to convert the video and audio from old videotapes into high-quality footage that users can easily play and share on modern devices.
[0046] The processing flow will be explained below.
[0047] Step 1:
[0048] Users convert old videotapes (VHS, miniDV, etc.) into digital data using dedicated digitizing equipment. The digitized video data is then stored on the user's device.
[0049] Step 2:
[0050] The terminal provides an interface for uploading the digitized video data to the server through a dedicated website or application. The user selects a file using the terminal interface and clicks the "upload" button.
[0051] Step 3:
[0052] The server receives the video data uploaded by the user and temporarily stores it. When storing it, it analyzes and extracts the metadata of the video data (resolution, frame rate, audio format, etc.).
[0053] Step 4:
[0054] The server then calls the generative AI model, which uses deep learning techniques to enhance the image quality of each frame, remove noise and blur, and perform color correction.
[0055] Step 5:
[0056] The server uses a generative AI model to process the extracted audio track to improve its quality, specifically by removing noise, canceling echoes, and adjusting pitch and clearing to improve sound quality.
[0057] Step 6:
[0058] The server recombines the high-quality video and audio data to generate a new video file, which is then reformatted to the user's desired format (e.g., DVD or MP4).
[0059] Step 7:
[0060] The server stores the generated reformatted video file and makes it accessible to users via a download link, which is communicated to users via email and / or their dashboard.
[0061] Step 8:
[0062] The user can then download the converted video file to their device using the provided download link, then burn it to a DVD or play it on their smartphone.
[0063] Step 9:
[0064] The device provides a communication application for sharing downloaded video files with family and friends. The user can use this application to upload video files to a cloud service or send them directly to family and friends.
[0065] Example 1
[0066] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0067] Many households currently hold old analog recording media (such as VHS tapes and miniDV) that have been recorded in the past, but they face the problem of being difficult to play back on modern devices. Another issue is the deterioration of these analog recording media, which leads to a decline in the quality of the video and audio. There is a need to solve these problems and convert valuable video assets from the past into high-quality digital data that can be easily played back and shared on modern devices.
[0068] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0069] In this invention, the server includes means for receiving digitized video data from old analog recording media, means for converting the data into digital files using dedicated digitizing equipment, means for improving the image quality of the digitized video data using a generative AI model, means for improving the audio quality of the digitized video data using a generative AI model, means for converting the improved image and audio quality video data into a format specified by a user, and means for providing the reformatted video data to a user. This allows the video and audio from old analog recording media to be converted into high-quality digital data that can be easily played and shared on modern devices.
[0070] An "analog recording medium" is a medium that records video and audio as analog signals, such as VHS tapes and miniDV.
[0071] "Digitalization" means converting video and audio recorded as analog signals into digital signals.
[0072] A "generative AI model" is an artificial intelligence model that uses machine learning and deep learning and is used to improve the quality of video and audio.
[0073] "High image quality" refers to processing to improve the image quality of video, and includes techniques such as improving resolution, removing noise, and correcting colors.
[0074] "High-quality sound" refers to processing to improve the quality of audio, and includes techniques such as noise removal and acoustic clearing.
[0075] "Formatting" refers to converting digital data into a format that can be played on a specific playback device or software, examples of which include MP4 and DVD ISO files.
[0076] A "download link" is a link that allows a user to obtain digital data via the Internet.
[0077] A "digital file" is a data file recorded as a digital signal, examples of which include video files and audio files.
[0078] "Specialized digitizing equipment" means equipment designed to digitize analog recording media, examples of which include video capture devices.
[0079] "Reformatting" means converting processed digital data into a different format.
[0080] The present invention is a system for digitizing video data from old analog recording media, converting it into high-quality image and sound using a generative AI model, and dubbing it into a format that can be played and shared on modern devices. The main aspects of implementing the present invention are described in detail below.
[0081] Video data digitization and uploading
[0082] Users use specialized digitizing equipment (e.g., video capture devices) to convert old analog recording media (e.g., VHS tapes, miniDV) into digital files. They save these digital files to their own devices and then access a specific website or application to upload the saved digital files to a server. An upload interface is provided, which users use to select files and start uploading.
[0083] Example: A user uses a home video capture device to digitize their VHS tapes and save them on their computer. Then, they access a dedicated website, select the saved video files, and upload them.
[0084] High-quality image and sound processing using generative AI models
[0085] The server receives and temporarily stores the uploaded video file. It then processes the video data using a generative AI model (e.g., Super Resolution GAN, Noise2Noise). The generative AI model performs the following tasks:
[0086] 1. Increase the resolution of each frame (using Super Resolution technology).
[0087] 2. Remove noise and blur (using De-Noising technology).
[0088] 3. Perform color correction (using Color Stability technology).
[0089] 4. The audio track is also subjected to noise removal and clearing to improve sound quality (using Audio Enhancement technology).
[0090] Example: A generative AI model upscales an original 480p home video to 1080p and produces a noise-reduced video. The audio track is processed to remove background noise and improve audio clarity.
[0091] Reformatting and providing data
[0092] Once processing is complete, the server uses software such as ffmpeg or HandBrake to convert the high-quality video file into the format of the user's choice (e.g., DVD ISO file, MP4 format). The converted data is then stored back on the server and a download link is generated. This link is then posted to the user for easy access.
[0093] User download and sharing of data
[0094] Users can download the converted video data using the provided link. They can then burn the downloaded data to a disc using home DVD burning software (e.g., Nero Burning ROM) or play it on their smartphone. They can also share the data with family and friends using cloud services or communication applications.
[0095] For example, a user downloads an ISO file provided by a server and burns it to a DVD disc using home DVD burning software. Then, the user plays the DVD on a home DVD player and watches it with his or her family. Alternatively, the user uploads the converted MP4 file to cloud storage and shares it with family and friends using a specific communication application.
[0096] Prompt Sentence Examples
[0097] "I have digitized some old VHS tapes. I would like to convert these files to high-quality MP4 format that can be played on my smartphone. Please use your generative AI model to convert them."
[0098] Thus, the present invention is a system that uses generative AI to convert video and audio from old analog recording media into high-quality content that users can easily play and share on modern devices.
[0099] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0100] Step 1: User digitizes video data
[0101] Users convert old analog recording media (e.g., VHS tapes, miniDV) into digital files using dedicated digitizing equipment (e.g., video capture devices). They start software on a device connected to the digitizing equipment, press the play button, and the analog video and audio are converted into digital signals, which are then saved on the device as digital files (e.g., .mp4, .avi).
[0102] Input: Analog recording media
[0103] Output: Digitized video data file
[0104] What it does: A user inserts a VHS tape into a video capture device, runs the software, and begins recording. Once the recording is complete, it is saved as a digital file on the device.
[0105] Step 2: User uploads digital data
[0106] The user accesses a specific website or application on his / her own terminal to upload the digital data saved in the previous step to the server, selects the file from the upload interface of the website or application, and clicks the upload button to start the file transfer.
[0107] Input: Digitized video data file
[0108] Output: Video data file saved on the server
[0109] Specific operation: The user opens a browser, accesses the specified URL, logs in, selects the digitized video file in the file upload interface, and clicks the upload button.
[0110] Step 3: High-quality image and sound processing by the server
[0111] The server receives the uploaded video file, temporarily stores it, and then uses a generative AI model (e.g., Super Resolution GAN, Noise2Noise) to enhance the image and sound quality of the video data.
[0112] 1. Apply Super Resolution technology to improve the resolution of each frame.
[0113] 2. De-Noising technology is applied to remove noise and blur.
[0114] 3. Color Stability technology is applied to perform color correction.
[0115] 4. Apply Audio Enhancement technology to denoise and clear the audio track.
[0116] Input: Uploaded video file
[0117] Output: High-quality video and audio files
[0118] What happens: The server inputs a prompt into the generative AI model, requesting, "Please improve the image and sound quality of this video." The generative AI model then begins processing, optimizing each frame and audio track.
[0119] Step 4: Reformatting Data by the Server
[0120] The processed video file is converted to the format specified by the user. The server uses tools such as ffmpeg or HandBrake to reformat the file and saves the newly generated file on the server.
[0121] Input: High-quality video files
[0122] Output: Video file in the specified format
[0123] Specific operation: The server runs the ffmpeg command to convert the video file to MP4 format or a DVD ISO file, and then saves the converted file.
[0124] Step 5: Generate and notify the download link
[0125] The server generates a download link for the reformatted video data and notifies the user via email or web notification.
[0126] Input: Reformatted video file
[0127] Output: Download link
[0128] What happens: The server generates a download link and sends it to the user's email address, or displays a notification on the web interface.
[0129] Step 6: Users download and share their data
[0130] The user can then use the provided link to download the converted video data, which can then be burned to a disc using home DVD burning software, played on a smartphone, or shared via cloud services or communication applications.
[0131] Input: Download link
[0132] Output: High-quality video data stored on home devices or cloud storage
[0133] What it does: A user clicks the download link and saves the file to their device. They then use DVD burning software to burn the ISO file to a disc and play it on a DVD player. Alternatively, they can play the downloaded MP4 file on their smartphone. They can also upload it to cloud storage and share it with family and friends.
[0134] (Application example 1)
[0135] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0136] One challenge is the difficulty of playing back memorable footage stored on old videotapes with high image and sound quality on modern devices. Additionally, there are limited ways for users to easily digitize these videos and obtain them in the format of their choice. Furthermore, traditional methods for remastering videos require specialized knowledge and software, making them inaccessible to the average user.
[0137] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0138] In this invention, the server includes means for receiving video data digitized from old videotapes, means for improving the image quality of the digitized video data using a generative AI model, means for improving the audio quality of the digitized video data using a generative AI model, means for reformatting the video data with improved image quality and audio quality, means for providing the reformatted video data to a user, and means for improving the image quality and audio quality of the video data on the cloud server and making the video data converted into a format selected by the user available for download. This allows users to easily convert old video footage into high quality and obtain it in a format playable on modern devices.
[0139] "Old videotapes" are analog storage media such as VHS and miniDV, and are tape devices used to store video and audio data.
[0140] "Digital video data" refers to video and audio data that has been converted from analog videotape and stored as a digital data file on a computer or other device.
[0141] A "generative AI model" is a type of artificial intelligence based on deep learning technology that generates new data based on input data. It has the potential to improve the quality of video and audio.
[0142] "High-definition" refers to the process of improving the image quality by increasing the resolution of existing images, removing noise and blur, and performing color correction.
[0143] "High-quality sound enhancement" is an acoustic process that removes background noise from the audio track and improves intelligibility and clarity of the sound.
[0144] "Reformatting" refers to converting the processed data into a different file format to optimize it for a specific playback device or purpose, such as MP4, MKV, or DVD.
[0145] A "cloud server" is a remote server available via the Internet, a computer system that provides large-volume data processing and storage.
[0146] A "format" refers to the data storage format or structure, and is a rule for storing digital data in a specific way so that it can be played or used.
[0147] This invention is a system that uses a generative AI model to convert video data digitized from old videotapes into high-quality image and sound, and dubs it into a format playable on modern devices. The system is designed to allow users to easily upload video data and download high-quality video data processed by the generative AI model.
[0148] Uploading video data
[0149] First, a user digitizes old videotapes (e.g., VHS or miniDV) using a home video capture device and saves them as digital files on their device. Then, the user opens a dedicated website or application and uploads the digitized video files to a server. The user selects the files using an intuitive interface and starts the upload. For example, a user can digitize a VHS tape at home, save it on their computer, and then access a dedicated website to select and upload the video files.
[0150] High-quality image and sound processing
[0151] The server receives and temporarily stores the uploaded video file. It then processes the video data using a generative AI model. The generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It can also denoise and clear the audio track, resulting in a higher quality sound. For example, it can upscale an original 480p home video to 1080p and remove background noise from the audio track for a clearer sound.
[0152] Reformatting and providing data
[0153] After the generative AI model completes the high-quality image and sound processing, the server reformats the processed video data. The user can then select the desired output format (e.g., DVD or MP4) and download the converted video data. The user obtains the video data through the generated download link. The user downloads the provided ISO file, burns it to a DVD using home DVD burning software, and plays it on a home DVD player. The converted MP4 file can also be uploaded to cloud storage and shared with family and friends through certain communication applications. An example prompt is, "Upload your old VHS home video and upscale it to high-quality 1080p with clear audio. This video will be available for download in MP4 format."
[0154] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0155] Step 1:
[0156] A user digitizes old videotapes.
[0157] Input: Video tape, video capture device
[0158] Output: Digital file
[0159] What it does: A user uses a home video capture device to convert video and audio data from an old videotape into digital files and save them on their device.
[0160] Step 2:
[0161] A user uploads a digitized video file to a server.
[0162] Input: Digital file, website upload interface
[0163] Output: Video file temporarily saved on the server
[0164] How it works: A user accesses a dedicated website or application, clicks the upload button, selects a digitized video file, and uploads it to the server.
[0165] Step 3:
[0166] The server processes the video data to improve its image and sound quality.
[0167] Input: Digital files, generative AI models
[0168] Output: High-quality video data
[0169] How it works: The server processes the uploaded video file using a generative AI model. Specifically, it improves the video resolution, removes noise and blur, and performs color correction. It also denoises and clears the audio track. The result is video and audio of higher quality than the original.
[0170] Step 4:
[0171] The server reformats the processed video data.
[0172] Input: High-quality video data, user-selectable format
[0173] Output: Video file in the specified format
[0174] What happens: The server converts the processed video data into the format selected by the user (e.g. MP4, DVD, etc.) using the appropriate encoding software.
[0175] Step 5:
[0176] The server generates a download link for the video data and provides it to the user.
[0177] Input: Format converted video file
[0178] Output: Download link
[0179] Specific operation: The server hosts the converted video file and generates a download link that users can access. The server notifies users of this link so that they can easily obtain the data.
[0180] Step 6:
[0181] Users can download and use high-quality video files.
[0182] Input: Download link
[0183] Output: Locally saved video file
[0184] Specific operation: Users can click the download link provided to download the high-quality video file to their device and play or share it.
[0185] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0186] The present invention is a system that uses generation AI to convert video data digitized from old videotapes into high-quality image and sound, combines it with an emotion engine that recognizes the user's emotions, processes and reformats the data according to the user's emotions, and finally dubs the data into a format that can be played on DVDs, smartphones, etc. The main modes for implementing the present invention are described in detail below.
[0187] Uploading video data
[0188] Users convert old videotapes (e.g., VHS or miniDV) into digital files using a dedicated digitizing device and save them on their device. Next, they open a specific website or application on their device to upload the saved digital data to a server. The web page provides an interface for selecting and uploading files.
[0189] Example: A user uses a home video capture device to digitize their VHS tapes and save them on their computer. Then, they access a dedicated website, select the saved video files, and upload them.
[0190] High-quality image and sound processing
[0191] The server receives and temporarily stores the uploaded video file. It then processes the received data using a generative AI model. The generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs noise reduction and clearing on the audio track to improve sound quality.
[0192] Example: A generative AI model upscales an original 480p home video to 1080p and produces a noise-reduced video. The audio track is processed to remove background noise and improve audio clarity.
[0193] User Emotion Recognition
[0194] The server analyzes the high-quality video data and invokes an emotion engine to recognize the user's emotions. The emotion engine analyzes the facial expressions and tone of voice of the people in the video data to detect emotions such as smiling or crying. Based on this emotion, filters and effects are added to specific parts of the video.
[0195] Example: An emotion engine analyzes a scene in a home video where a child blows out a birthday cake, and if it detects a happy expression, it adds a special effect to the scene.
[0196] Reformatting and providing data
[0197] Once the processing is complete, the server converts the high-quality video file to the format of the user's choice (e.g. DVD, MP4, etc.). The converted data is then re-saved and a download link is generated. This link is then communicated to the user for easy access.
[0198] The user can then use the provided link to download the converted video data to their device. They can then burn the downloaded file to a DVD, play it on their smartphone, or share the video with family and friends using a communication app or cloud service for data sharing.
[0199] For example, a user downloads an ISO file provided by a server and burns it to a DVD disc using home DVD burning software. Then, the user plays the DVD on a home DVD player and watches it with his or her family. Alternatively, the user uploads the converted MP4 file to cloud storage and shares it with family and friends using a specific communication application.
[0200] Thus, the present invention is a system that uses generative AI and an emotion engine to convert the video and audio from old videotapes into high-quality footage, apply specific processing based on the user's emotions, and make it easy to play and share on modern devices.
[0201] The processing flow will be explained below.
[0202] Step 1:
[0203] Users convert old videotapes (VHS, miniDV, etc.) into digital data using dedicated digitizing equipment. The digitized video data is then stored on the user's device.
[0204] Step 2:
[0205] The terminal provides an interface for uploading the digitized video data to the server through a dedicated website or application. The user selects a file using the terminal interface and clicks the "upload" button.
[0206] Step 3:
[0207] The server receives the video data uploaded by the user and temporarily stores it. When storing it, it analyzes and extracts the metadata of the video data (resolution, frame rate, audio format, etc.).
[0208] Step 4:
[0209] The server then calls the generative AI model, which uses deep learning techniques to enhance the image quality of each frame, remove noise and blur, and perform color correction.
[0210] Step 5:
[0211] The server uses a generative AI model to process the extracted audio track to improve its quality, specifically by removing noise, canceling echoes, and adjusting pitch and clearing to improve sound quality.
[0212] Step 6:
[0213] The server recombines the high-quality video and audio data to generate a new video file, which is then reformatted to the format specified by the user (e.g., DVD format, MP4 format).
[0214] Step 7:
[0215] The server invokes the emotion engine to analyze the facial expressions and voices of people in the video data, and the emotion engine detects specific emotions, such as smiling or crying.
[0216] Step 8:
[0217] The server processes the video by adding filters and effects to specific parts of the video based on the emotions detected by the emotion engine. For example, a special effect is applied to scenes where the emotion of joy is detected.
[0218] Step 9:
[0219] The server reformats the emotion-processed, high-quality video data into a format specified by the user, and generates and saves the final video file.
[0220] Step 10:
[0221] The server stores the generated reformatted video file and makes it accessible to users via a download link, which is communicated to users via email and / or their dashboard.
[0222] Step 11:
[0223] The user can then download the converted video file to their device using the provided download link, then burn it to a DVD or play it on their smartphone.
[0224] Step 12:
[0225] The device provides a communication application for sharing downloaded video files with family and friends. The user can use this application to upload video files to a cloud service or send them directly to family and friends.
[0226] Example 2
[0227] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0228] Due to aging and differences in formats, it is difficult to play the video and audio from old videotapes on modern devices. Furthermore, these videos suffer from noise and low resolution, resulting in a poor viewing experience. Furthermore, there is a lack of technology to recognize user emotions and enrich the visual expression. Therefore, current technology is required to improve the image and sound quality of old videotapes, convert them into formats playable on modern devices, and even add effects that respond to user emotions.
[0229] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for receiving video data digitized from an old video tape, a means for improving the image quality of the digitized video data using a generative AI model, and a means for improving the sound quality of the audio of the digitized video data using a generative AI model. This makes it possible to analyze the video data with high image quality and sound quality and recognize the user's emotions. In addition, by including a means for adding filters and effects to the video data based on the recognized emotions, the visual expression can be enriched.
[0230] "Digitization" is the process of converting analog data into digital data.
[0231] A "generative AI model" is an algorithm that uses artificial intelligence to generate or transform data for a specific task using techniques such as deep learning.
[0232] "High-definition" refers to the process of improving the resolution and image quality of video data.
[0233] "High quality audio" is the process of improving the quality of audio data and removing noise.
[0234] "Emotion recognition" is a technology that analyzes facial expressions and vocal tones of people contained in video data to detect specific emotional states (such as joy, anger, sadness, or happiness).
[0235] A "filter" is a function for applying a specific effect to video data.
[0236] "Effects" is a function that adds special visual effects to video data.
[0237] "Reformatting" is the process of converting video data into a different form or format.
[0238] "Providing" refers to making the generated data available to users.
[0239] "Server" means a central control unit for receiving, storing, processing, and providing video data.
[0240] "User" means the person or end user who utilizes the system to upload video data and download the final product.
[0241] This invention is a system that uses a generative AI model to convert video data digitized from old videotapes into high-quality image and sound, combines it with an emotion engine that recognizes the user's emotions, processes and reformats the data according to the user's emotions, and finally dubs it into a format that can be played on DVDs, smartphones, etc. Specific steps for implementing this invention are described below.
[0242] Video data digitization and uploading
[0243] A user converts old videotapes (e.g., VHS or miniDV) into digital files using dedicated digitizing equipment. This can be done using a home video capture device. The digitized video data is saved on the device. The user then opens a specific website or application on the device and uploads the saved video data to a server. The website provides an interface for selecting and uploading files, and the user selects and uploads the digital data.
[0244] High-quality image and sound processing
[0245] The server receives and temporarily stores video files uploaded by users. The server then launches a generative AI model to process the uploaded video data. This generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs similar noise reduction and sound quality clearing on the audio track to improve sound quality. For example, the generative AI model upscales a 480p home video to 1080p and generates a noise-free video. It also performs processing to remove background noise from the audio track and improve sound quality.
[0246] User Emotion Recognition
[0247] The server analyzes the high-quality video data and invokes an emotion engine to recognize the user's emotions. The emotion engine analyzes the facial expressions and tone of voice of the people in the video data to detect emotions such as smiling or crying. Based on these emotions, filters and effects are added to specific parts of the video. For example, if the emotion engine analyzes a scene in a home video where a child blows out a birthday cake and detects a happy expression, it adds a special effect to that scene.
[0248] Reformatting and providing data
[0249] The server converts the processed video file with high image and sound quality into the format specified by the user (e.g., DVD, MP4, etc.). The converted data is then saved again, and a download link is generated and sent to the user. The user then downloads the video data to their device via the link. The user can then burn the downloaded file to a DVD or play it on their smartphone. They can also share the video with family and friends using cloud services or communication applications. For example, a user can download an ISO file provided by the server and burn it to a DVD using their home DVD burning software. They can then play the DVD on a home DVD player and enjoy it with their family. They can also upload the converted MP4 file to cloud storage and share it with family and friends using a specific communication application.
[0250] Prompt Sentence Examples
[0251] 1. "What are the specific steps to digitize and upload my old VHS tapes?"
[0252] 2. "How can I use a generative AI model to convert uploaded videos into high-quality video and audio?"
[0253] 3. "Please explain how you can analyze video data, recognize user emotions, and add specific effects."
[0254] The above is a specific embodiment for implementing the present invention. The present invention is a system for converting old videotape digital data into high-quality data, processing it according to the user's emotions, and making it easy to play and share on modern devices.
[0255] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0256] Step 1: Digitize and upload your video data
[0257] A user converts old videotapes (e.g., VHS or miniDV) into digital files using dedicated digitizing equipment (e.g., a home video capture device). The converted digital video data is stored on the device. Next, the user opens a specific website or application on the device to upload the saved digital data to a server. Through the file upload interface, the user selects the video data and clicks the upload button. The server receives this video data and temporarily stores it. The input is a digitized video file, and the output is video data temporarily stored on the server.
[0258] Step 2: High-quality image and sound processing
[0259] The server retrieves the stored video data and launches a generative AI model. When the video data is input into the generative AI model, it uses deep learning technology to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs noise reduction and sound quality clearing on the audio track. As a specific example, the generative AI model upscales a 480p home video to 1080p and generates a noise-free video. At the same time, it removes background noise from the audio track and performs processing to clear the sound quality. The input is temporarily stored video data, and the output is video data with high image quality and sound quality.
[0260] Step 3: Recognizing user emotions
[0261] The server receives high-quality video data from the generative AI model and passes it to the emotion engine. When video data is input into the emotion engine, it analyzes the facial expressions and voice tones of the people in the video to detect emotions such as smiling or crying. Based on the emotions recognized by the emotion engine, filters and effects are added to specific scenes. For example, if the emotion engine analyzes a scene of a child blowing out a birthday cake and detects a joyful expression, it adds a special effect to that scene. The input is high-quality video data, and the output is video data with emotion recognition and effects applied.
[0262] Step 4: Reformat and provide data
[0263] Once processing is complete, the server converts the video data, which has undergone high-quality image and sound quality enhancement and emotion recognition processing, into the format specified by the user (e.g., DVD, MP4, etc.). The converted data is then saved back to the server, and a download link is generated. The user then downloads the converted video data to their device via this download link. For example, a user downloads an ISO file provided by the server and burns it to a DVD using their home DVD burning software. The user then plays the DVD on a home DVD player and watches it with their family. Alternatively, the user can upload the converted MP4 file to cloud storage and share it with family and friends using a specific communication application. The input is video data with emotion recognition and effects applied, and the output is video data converted to the user-specified format.
[0264] The above describes the specific operations and data flow for each processing step. This system allows us to convert old videotapes to high quality, add visual effects based on the user's emotions, and provide them in a format that can be played and shared on modern devices.
[0265] (Application example 2)
[0266] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0267] Video and audio stored on old videotapes are not suitable for playback on modern devices with high image and sound quality. Furthermore, digitized data from these videotapes suffers from degradation of the original image and sound quality, making it difficult to achieve visual and auditory satisfaction. Furthermore, the video recorded on home videotapes is filled with people's memories and emotions, and there is a demand for content that reflects those emotions. To solve these issues, not only is it necessary to improve the quality of video data, but it is also necessary to edit it in a way that reflects the user's emotions.
[0268] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0269] In this invention, the server includes means for receiving video data digitized from old videotapes, means for improving the image quality of the digitized video data using a generative AI model, means for improving the audio quality of the digitized video data using a generative AI model, means for reformatting the video data with improved image quality and audio quality, means for providing the reformatted video data to a user, means for recognizing the emotions of people in the video data using an emotion engine and adding effects according to the recognized emotions, and means for outputting the video data in a specified format. This makes it possible to convert data from old videotapes into high-quality image and audio and provide it as video that reflects the user's emotions.
[0270] "Old videotapes" are magnetic tape media that record video and audio in a non-digital analog format, such as VHS or miniDV.
[0271] "Digitalized video data" means video and audio data recorded on analog videotape that has been converted into a digital format that can be processed and stored on an electronic device.
[0272] A "generative AI model" is an artificial intelligence model designed based on deep learning technology, and is an algorithm that analyzes and processes input data to generate specific results.
[0273] "High-definition" refers to the process of processing video data that has been degraded by low resolution or noise using a generative AI model to convert it into high-resolution, clear video.
[0274] "High-quality sound" refers to processing audio data containing background noise and degraded sound using a generative AI model to convert it into clear, high-quality sound.
[0275] "Reformatting" refers to the process of converting digitized video data into a specific file format or designated media format.
[0276] An "emotion engine" is an artificial intelligence technology that analyzes facial expressions and tone of voice of people contained in video and audio data and recognizes their emotions.
[0277] "Effects" refers to the process of adding specific visual or auditory effects to video or audio.
[0278] "Video data metadata" refers to information related to video data, such as supplementary information such as the date and time of filming, the location of filming, and the characters appearing in the video.
[0279] The "specified format" refers to the particular file or media format selected by the user for the final output of the video data.
[0280] This invention is a system that uses generative AI to convert video data digitized from old videotapes into high-quality image and sound, then combines it with an emotion engine that recognizes the user's emotions to add specific effects to the video data, and finally provides it in a format specified by the user.
[0281] Overall system configuration
[0282] The system mainly includes the following components:
[0283] 1. Digitization tools: Video capture devices to generate digital data from old videotapes.
[0284] 2. Data receiving means: A terminal (e.g., a smartphone or PC) that provides an interface for uploading digitized video data to a server.
[0285] 3. Generative AI model: A deep learning model for processing digital video and audio data to produce high-quality images and sounds.
[0286] 4. Emotion Engine: An artificial intelligence engine to analyze the emotions of people in video data and add specific effects.
[0287] 5. Reformatting means: A device that converts video data with high image quality, high sound quality, and added emotional effects into a format specified by the user (e.g., MP4, DVD ISO).
[0288] 6. Data providing means: A server that generates a download link to provide the final video data to the user.
[0289] What the program does
[0290] 1. Upload video data:
[0291] The user starts a dedicated application on a device (such as a smartphone or PC) and uploads the digitized files of old videotapes to the server. At this stage, a file selection interface is provided.
[0292] 2. Processing by generative AI models:
[0293] The server temporarily stores the uploaded video data and uses a generative AI model to improve the video and audio quality. This model is based on deep learning technology using TensorFlow. Specifically, it improves the video resolution, removes noise and blur, and performs color correction. It also adjusts the audio track to remove background noise and ensure clear sound quality.
[0294] 3. Emotion engine processing:
[0295] The video data, which has been enhanced in quality by the generative AI model, is then analyzed using an emotion engine. The emotion engine uses OpenCV and other technologies to analyze the facial expressions and tone of voice of the people in the video, and adds specific effects depending on the emotions it recognizes. For example, in a video of a birthday party, a congratulatory effect can be added to scenes where a child is happy.
[0296] 4. Data Reformatting and Provision:
[0297] The final processed video data is converted into the format specified by the user (e.g. MP4 or DVD ISO). The server saves the converted video data in storage and notifies the user of a download link. Using this link, the user can download the video data and save it on their device.
[0298] Prompt Sentence Examples
[0299] "Convert your 480p home videos to high-quality 1080p footage, remove background noise and improve the clarity of your audio tracks."
[0300] As described above, the present invention is a system that converts digital data from old videotapes into high-quality data and adds effects that correspond to the user's emotions, thereby providing memorable videos in a more attractive form.
[0301] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0302] Step 1:
[0303] Users launch a dedicated application on their device (smartphone or PC), select the digitized files of old videotapes, and upload them to the server.
[0304] Input: Digitized video data file
[0305] Process: The user selects a file using the file selection interface and presses the "Upload" button.
[0306] Output: Video data is uploaded to the server and temporarily stored.
[0307] Step 2:
[0308] The server temporarily stores the uploaded video data and uses generative AI models to improve the video and audio quality.
[0309] Input: Uploaded video data
[0310] Processing: Generative AI models are used to improve video resolution, remove noise and blur, and perform color correction. Specifically, deep learning models using TensorFlow perform frame-by-frame processing. Audio tracks are also adjusted to remove background noise and ensure clarity.
[0311] Output: High-quality video data
[0312] Step 3:
[0313] The emotion engine analyzes high-quality video data and audio to recognize people's emotions.
[0314] Input: High-quality video data
[0315] Processing: An emotion engine (using OpenCV, etc.) is used to analyze facial expressions and vocal tones of people in the video to recognize specific emotions.
[0316] Output: Metadata about the recognized emotion
[0317] Step 4:
[0318] The server adds specific effects to the video data according to the emotions recognized by the emotion engine.
[0319] Input: High-quality video data and emotion recognition metadata
[0320] Processing: Add effects to specific scenes based on the recognized emotion. For example, add a blessing effect to a scene where the emotion is recognized as "joy."
[0321] Output: Video data with effects added
[0322] Step 5:
[0323] The server then converts the final processed video data into the format specified by the user and generates a download link.
[0324] Input: Video data with effects added
[0325] Processing: Convert the data to the desired format (e.g. MP4, DVD ISO) and save it to storage, using a reformatting tool to convert it to the appropriate file format, and generate a link for the user to download it.
[0326] Output: Video data in the specified format and a download link
[0327] Step 6:
[0328] The user uses the provided download link to download the final video data to their terminal.
[0329] Input: Download link
[0330] Processing: The user clicks on the notified link and downloads the video data.
[0331] Output: The final video data saved on your device
[0332] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0333] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0334] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0335] [Second embodiment]
[0336] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0337] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0338] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0339] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0340] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0341] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0342] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0343] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0344] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0345] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0346] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0347] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0348] The present invention is a system for converting video data digitized from old videotapes into high-quality image and sound using generation AI, and dubbing it into a format that can be played on DVDs, smartphones, etc. The main modes for implementing the present invention are described in detail below.
[0349] Uploading video data
[0350] Users convert old videotapes (e.g., VHS or miniDV) into digital files using a dedicated digitizing device and save them on their device. Next, they open a specific website or application on their device to upload the saved digital data to a server. The web page provides an interface for selecting and uploading files.
[0351] Example: A user uses a home video capture device to digitize their VHS tapes and save them on their computer. Then, they access a dedicated website, select the saved video files, and upload them.
[0352] High-quality image and sound processing
[0353] The server receives and temporarily stores the uploaded video file. It then processes the received data using a generative AI model. The generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs noise reduction and clearing on the audio track to improve sound quality.
[0354] Example: A generative AI model upscales an original 480p home video to 1080p and produces a noise-reduced video. The audio track is processed to remove background noise and improve audio clarity.
[0355] Reformatting and providing data
[0356] Once the processing is complete, the server converts the high-quality video file to the format of the user's choice (e.g. DVD, MP4, etc.). The converted data is then re-saved and a download link is generated. This link is then communicated to the user for easy access.
[0357] Users can download the converted video data using the provided link, burn it to a disc for playback on a home DVD player, play it on a smartphone, or share it with family and friends using a communication application or cloud service for data sharing.
[0358] For example, a user downloads an ISO file provided by a server and burns it to a DVD disc using home DVD burning software. Then, the user plays the DVD on a home DVD player and watches it with his or her family. Alternatively, the user uploads the converted MP4 file to cloud storage and shares it with family and friends using a specific communication application.
[0359] Thus, the present invention is a system that uses generative AI to convert the video and audio from old videotapes into high-quality footage that users can easily play and share on modern devices.
[0360] The processing flow will be explained below.
[0361] Step 1:
[0362] Users convert old videotapes (VHS, miniDV, etc.) into digital data using dedicated digitizing equipment. The digitized video data is then stored on the user's device.
[0363] Step 2:
[0364] The terminal provides an interface for uploading the digitized video data to the server through a dedicated website or application. The user selects a file using the terminal interface and clicks the "upload" button.
[0365] Step 3:
[0366] The server receives the video data uploaded by the user and temporarily stores it. When storing it, it analyzes and extracts the metadata of the video data (resolution, frame rate, audio format, etc.).
[0367] Step 4:
[0368] The server then calls the generative AI model, which uses deep learning techniques to enhance the image quality of each frame, remove noise and blur, and perform color correction.
[0369] Step 5:
[0370] The server uses a generative AI model to process the extracted audio track to improve its quality, specifically by removing noise, canceling echoes, and adjusting pitch and clearing to improve sound quality.
[0371] Step 6:
[0372] The server recombines the high-quality video and audio data to generate a new video file, which is then reformatted to the user's desired format (e.g., DVD or MP4).
[0373] Step 7:
[0374] The server stores the generated reformatted video file and makes it accessible to users via a download link, which is communicated to users via email and / or their dashboard.
[0375] Step 8:
[0376] The user can then download the converted video file to their device using the provided download link, then burn it to a DVD or play it on their smartphone.
[0377] Step 9:
[0378] The device provides a communication application for sharing downloaded video files with family and friends. The user can use this application to upload video files to a cloud service or send them directly to family and friends.
[0379] Example 1
[0380] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0381] Many households currently hold old analog recording media (such as VHS tapes and miniDV) that have been recorded in the past, but they face the problem of being difficult to play back on modern devices. Another issue is the deterioration of these analog recording media, which leads to a decline in the quality of the video and audio. There is a need to solve these problems and convert valuable video assets from the past into high-quality digital data that can be easily played back and shared on modern devices.
[0382] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0383] In this invention, the server includes means for receiving digitized video data from old analog recording media, means for converting the data into digital files using dedicated digitizing equipment, means for improving the image quality of the digitized video data using a generative AI model, means for improving the audio quality of the digitized video data using a generative AI model, means for converting the improved image and audio quality video data into a format specified by a user, and means for providing the reformatted video data to a user. This allows the video and audio from old analog recording media to be converted into high-quality digital data that can be easily played and shared on modern devices.
[0384] An "analog recording medium" is a medium that records video and audio as analog signals, such as VHS tapes and miniDV.
[0385] "Digitalization" means converting video and audio recorded as analog signals into digital signals.
[0386] A "generative AI model" is an artificial intelligence model that uses machine learning and deep learning and is used to improve the quality of video and audio.
[0387] "High image quality" refers to processing to improve the image quality of video, and includes techniques such as improving resolution, removing noise, and correcting colors.
[0388] "High-quality sound" refers to processing to improve the quality of audio, and includes techniques such as noise removal and acoustic clearing.
[0389] "Formatting" refers to converting digital data into a format that can be played on a specific playback device or software, examples of which include MP4 and DVD ISO files.
[0390] A "download link" is a link that allows a user to obtain digital data via the Internet.
[0391] A "digital file" is a data file recorded as a digital signal, examples of which include video files and audio files.
[0392] "Specialized digitizing equipment" means equipment designed to digitize analog recording media, examples of which include video capture devices.
[0393] "Reformatting" means converting processed digital data into a different format.
[0394] The present invention is a system for digitizing video data from old analog recording media, converting it into high-quality image and sound using a generative AI model, and dubbing it into a format that can be played and shared on modern devices. The main aspects of implementing the present invention are described in detail below.
[0395] Video data digitization and uploading
[0396] Users use specialized digitizing equipment (e.g., video capture devices) to convert old analog recording media (e.g., VHS tapes, miniDV) into digital files. They save these digital files to their own devices and then access a specific website or application to upload the saved digital files to a server. An upload interface is provided, which users use to select files and start uploading.
[0397] Example: A user uses a home video capture device to digitize their VHS tapes and save them on their computer. Then, they access a dedicated website, select the saved video files, and upload them.
[0398] High-quality image and sound processing using generative AI models
[0399] The server receives and temporarily stores the uploaded video file. It then processes the video data using a generative AI model (e.g., Super Resolution GAN, Noise2Noise). The generative AI model performs the following tasks:
[0400] 1. Increase the resolution of each frame (using Super Resolution technology).
[0401] 2. Remove noise and blur (using De-Noising technology).
[0402] 3. Perform color correction (using Color Stability technology).
[0403] 4. The audio track is also subjected to noise removal and clearing to improve sound quality (using Audio Enhancement technology).
[0404] Example: A generative AI model upscales an original 480p home video to 1080p and produces a noise-reduced video. The audio track is processed to remove background noise and improve audio clarity.
[0405] Reformatting and providing data
[0406] Once processing is complete, the server uses software such as ffmpeg or HandBrake to convert the high-quality video file into the format of the user's choice (e.g., DVD ISO file, MP4 format). The converted data is then stored back on the server and a download link is generated. This link is then posted to the user for easy access.
[0407] User download and sharing of data
[0408] Users can download the converted video data using the provided link. They can then burn the downloaded data to a disc using home DVD burning software (e.g., Nero Burning ROM) or play it on their smartphone. They can also share the data with family and friends using cloud services or communication applications.
[0409] For example, a user downloads an ISO file provided by a server and burns it to a DVD disc using home DVD burning software. Then, the user plays the DVD on a home DVD player and watches it with his or her family. Alternatively, the user uploads the converted MP4 file to cloud storage and shares it with family and friends using a specific communication application.
[0410] Prompt Sentence Examples
[0411] "I have digitized some old VHS tapes. I would like to convert these files to high-quality MP4 format that can be played on my smartphone. Please use your generative AI model to convert them."
[0412] Thus, the present invention is a system that uses generative AI to convert video and audio from old analog recording media into high-quality content that users can easily play and share on modern devices.
[0413] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0414] Step 1: User digitizes video data
[0415] Users convert old analog recording media (e.g., VHS tapes, miniDV) into digital files using dedicated digitizing equipment (e.g., video capture devices). They start software on a device connected to the digitizing equipment, press the play button, and the analog video and audio are converted into digital signals, which are then saved on the device as digital files (e.g., .mp4, .avi).
[0416] Input: Analog recording media
[0417] Output: Digitized video data file
[0418] What it does: A user inserts a VHS tape into a video capture device, runs the software, and begins recording. Once the recording is complete, it is saved as a digital file on the device.
[0419] Step 2: User uploads digital data
[0420] The user accesses a specific website or application on his / her own terminal to upload the digital data saved in the previous step to the server, selects the file from the upload interface of the website or application, and clicks the upload button to start the file transfer.
[0421] Input: Digitized video data file
[0422] Output: Video data file saved on the server
[0423] Specific operation: The user opens a browser, accesses the specified URL, logs in, selects the digitized video file in the file upload interface, and clicks the upload button.
[0424] Step 3: High-quality image and sound processing by the server
[0425] The server receives the uploaded video file, temporarily stores it, and then uses a generative AI model (e.g., Super Resolution GAN, Noise2Noise) to enhance the image and sound quality of the video data.
[0426] 1. Apply Super Resolution technology to improve the resolution of each frame.
[0427] 2. De-Noising technology is applied to remove noise and blur.
[0428] 3. Color Stability technology is applied to perform color correction.
[0429] 4. Apply Audio Enhancement technology to denoise and clear the audio track.
[0430] Input: Uploaded video file
[0431] Output: High-quality video and audio files
[0432] What happens: The server inputs a prompt into the generative AI model, requesting, "Please improve the image and sound quality of this video." The generative AI model then begins processing, optimizing each frame and audio track.
[0433] Step 4: Reformatting Data by the Server
[0434] The processed video file is converted to the format specified by the user. The server uses tools such as ffmpeg or HandBrake to reformat the file and saves the newly generated file on the server.
[0435] Input: High-quality video files
[0436] Output: Video file in the specified format
[0437] Specific operation: The server runs the ffmpeg command to convert the video file to MP4 format or a DVD ISO file, and then saves the converted file.
[0438] Step 5: Generate and notify the download link
[0439] The server generates a download link for the reformatted video data and notifies the user via email or web notification.
[0440] Input: Reformatted video file
[0441] Output: Download link
[0442] What happens: The server generates a download link and sends it to the user's email address, or displays a notification on the web interface.
[0443] Step 6: Users download and share their data
[0444] The user can then use the provided link to download the converted video data, which can then be burned to a disc using home DVD burning software, played on a smartphone, or shared via cloud services or communication applications.
[0445] Input: Download link
[0446] Output: High-quality video data stored on home devices or cloud storage
[0447] What it does: A user clicks the download link and saves the file to their device. They then use DVD burning software to burn the ISO file to a disc and play it on a DVD player. Alternatively, they can play the downloaded MP4 file on their smartphone. They can also upload it to cloud storage and share it with family and friends.
[0448] (Application example 1)
[0449] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0450] One challenge is the difficulty of playing back memorable footage stored on old videotapes with high image and sound quality on modern devices. Additionally, there are limited ways for users to easily digitize these videos and obtain them in the format of their choice. Furthermore, traditional methods for remastering videos require specialized knowledge and software, making them inaccessible to the average user.
[0451] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0452] In this invention, the server includes means for receiving video data digitized from old videotapes, means for improving the image quality of the digitized video data using a generative AI model, means for improving the audio quality of the digitized video data using a generative AI model, means for reformatting the video data with improved image quality and audio quality, means for providing the reformatted video data to a user, and means for improving the image quality and audio quality of the video data on the cloud server and making the video data converted into a format selected by the user available for download. This allows users to easily convert old video footage into high quality and obtain it in a format playable on modern devices.
[0453] "Old videotapes" are analog storage media such as VHS and miniDV, and are tape devices used to store video and audio data.
[0454] "Digital video data" refers to video and audio data that has been converted from analog videotape and stored as a digital data file on a computer or other device.
[0455] A "generative AI model" is a type of artificial intelligence based on deep learning technology that generates new data based on input data. It has the potential to improve the quality of video and audio.
[0456] "High-definition" refers to the process of improving the image quality by increasing the resolution of existing images, removing noise and blur, and performing color correction.
[0457] "High-quality sound enhancement" is an acoustic process that removes background noise from the audio track and improves intelligibility and clarity of the sound.
[0458] "Reformatting" refers to converting the processed data into a different file format to optimize it for a specific playback device or purpose, such as MP4, MKV, or DVD.
[0459] A "cloud server" is a remote server available via the Internet, a computer system that provides large-volume data processing and storage.
[0460] A "format" refers to the data storage format or structure, and is a rule for storing digital data in a specific way so that it can be played or used.
[0461] This invention is a system that uses a generative AI model to convert video data digitized from old videotapes into high-quality image and sound, and dubs it into a format playable on modern devices. The system is designed to allow users to easily upload video data and download high-quality video data processed by the generative AI model.
[0462] Uploading video data
[0463] First, a user digitizes old videotapes (e.g., VHS or miniDV) using a home video capture device and saves them as digital files on their device. Then, the user opens a dedicated website or application and uploads the digitized video files to a server. The user selects the files using an intuitive interface and starts the upload. For example, a user can digitize a VHS tape at home, save it on their computer, and then access a dedicated website to select and upload the video files.
[0464] High-quality image and sound processing
[0465] The server receives and temporarily stores the uploaded video file. It then processes the video data using a generative AI model. The generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It can also denoise and clear the audio track, resulting in a higher quality sound. For example, it can upscale an original 480p home video to 1080p and remove background noise from the audio track for a clearer sound.
[0466] Reformatting and providing data
[0467] After the generative AI model completes the high-quality image and sound processing, the server reformats the processed video data. The user can then select the desired output format (e.g., DVD or MP4) and download the converted video data. The user obtains the video data through the generated download link. The user downloads the provided ISO file, burns it to a DVD using home DVD burning software, and plays it on a home DVD player. The converted MP4 file can also be uploaded to cloud storage and shared with family and friends through certain communication applications. An example prompt is, "Upload your old VHS home video and upscale it to high-quality 1080p with clear audio. This video will be available for download in MP4 format."
[0468] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0469] Step 1:
[0470] A user digitizes old videotapes.
[0471] Input: Video tape, video capture device
[0472] Output: Digital file
[0473] What it does: A user uses a home video capture device to convert video and audio data from an old videotape into digital files and save them on their device.
[0474] Step 2:
[0475] A user uploads a digitized video file to a server.
[0476] Input: Digital file, website upload interface
[0477] Output: Video file temporarily saved on the server
[0478] How it works: A user accesses a dedicated website or application, clicks the upload button, selects a digitized video file, and uploads it to the server.
[0479] Step 3:
[0480] The server processes the video data to improve its image and sound quality.
[0481] Input: Digital files, generative AI models
[0482] Output: High-quality video data
[0483] How it works: The server processes the uploaded video file using a generative AI model. Specifically, it improves the video resolution, removes noise and blur, and performs color correction. It also denoises and clears the audio track. The result is video and audio of higher quality than the original.
[0484] Step 4:
[0485] The server reformats the processed video data.
[0486] Input: High-quality video data, user-selectable format
[0487] Output: Video file in the specified format
[0488] What happens: The server converts the processed video data into the format selected by the user (e.g. MP4, DVD, etc.) using the appropriate encoding software.
[0489] Step 5:
[0490] The server generates a download link for the video data and provides it to the user.
[0491] Input: Format converted video file
[0492] Output: Download link
[0493] Specific operation: The server hosts the converted video file and generates a download link that users can access. The server notifies users of this link so that they can easily obtain the data.
[0494] Step 6:
[0495] Users can download and use high-quality video files.
[0496] Input: Download link
[0497] Output: Locally saved video file
[0498] Specific operation: Users can click the download link provided to download the high-quality video file to their device and play or share it.
[0499] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0500] The present invention is a system that uses generation AI to convert video data digitized from old videotapes into high-quality image and sound, combines it with an emotion engine that recognizes the user's emotions, processes and reformats the data according to the user's emotions, and finally dubs the data into a format that can be played on DVDs, smartphones, etc. The main modes for implementing the present invention are described in detail below.
[0501] Uploading video data
[0502] Users convert old videotapes (e.g., VHS or miniDV) into digital files using a dedicated digitizing device and save them on their device. Next, they open a specific website or application on their device to upload the saved digital data to a server. The web page provides an interface for selecting and uploading files.
[0503] Example: A user uses a home video capture device to digitize their VHS tapes and save them on their computer. Then, they access a dedicated website, select the saved video files, and upload them.
[0504] High-quality image and sound processing
[0505] The server receives and temporarily stores the uploaded video file. It then processes the received data using a generative AI model. The generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs noise reduction and clearing on the audio track to improve sound quality.
[0506] Example: A generative AI model upscales an original 480p home video to 1080p and produces a noise-reduced video. The audio track is processed to remove background noise and improve audio clarity.
[0507] User Emotion Recognition
[0508] The server analyzes the high-quality video data and invokes an emotion engine to recognize the user's emotions. The emotion engine analyzes the facial expressions and tone of voice of the people in the video data to detect emotions such as smiling or crying. Based on this emotion, filters and effects are added to specific parts of the video.
[0509] Example: An emotion engine analyzes a scene in a home video where a child blows out a birthday cake, and if it detects a happy expression, it adds a special effect to the scene.
[0510] Reformatting and providing data
[0511] Once the processing is complete, the server converts the high-quality video file to the format of the user's choice (e.g. DVD, MP4, etc.). The converted data is then re-saved and a download link is generated. This link is then communicated to the user for easy access.
[0512] The user can then use the provided link to download the converted video data to their device. They can then burn the downloaded file to a DVD, play it on their smartphone, or share the video with family and friends using a communication app or cloud service for data sharing.
[0513] For example, a user downloads an ISO file provided by a server and burns it to a DVD disc using home DVD burning software. Then, the user plays the DVD on a home DVD player and watches it with his or her family. Alternatively, the user uploads the converted MP4 file to cloud storage and shares it with family and friends using a specific communication application.
[0514] Thus, the present invention is a system that uses generative AI and an emotion engine to convert the video and audio from old videotapes into high-quality footage, apply specific processing based on the user's emotions, and make it easy to play and share on modern devices.
[0515] The processing flow will be explained below.
[0516] Step 1:
[0517] Users convert old videotapes (VHS, miniDV, etc.) into digital data using dedicated digitizing equipment. The digitized video data is then stored on the user's device.
[0518] Step 2:
[0519] The terminal provides an interface for uploading the digitized video data to the server through a dedicated website or application. The user selects a file using the terminal interface and clicks the "upload" button.
[0520] Step 3:
[0521] The server receives the video data uploaded by the user and temporarily stores it. When storing it, it analyzes and extracts the metadata of the video data (resolution, frame rate, audio format, etc.).
[0522] Step 4:
[0523] The server then calls the generative AI model, which uses deep learning techniques to enhance the image quality of each frame, remove noise and blur, and perform color correction.
[0524] Step 5:
[0525] The server uses a generative AI model to process the extracted audio track to improve its quality, specifically by removing noise, canceling echoes, and adjusting pitch and clearing to improve sound quality.
[0526] Step 6:
[0527] The server recombines the high-quality video and audio data to generate a new video file, which is then reformatted to the format specified by the user (e.g., DVD format, MP4 format).
[0528] Step 7:
[0529] The server invokes the emotion engine to analyze the facial expressions and voices of people in the video data, and the emotion engine detects specific emotions, such as smiling or crying.
[0530] Step 8:
[0531] The server processes the video by adding filters and effects to specific parts of the video based on the emotions detected by the emotion engine. For example, a special effect is applied to scenes where the emotion of joy is detected.
[0532] Step 9:
[0533] The server reformats the emotion-processed, high-quality video data into a format specified by the user, and generates and saves the final video file.
[0534] Step 10:
[0535] The server stores the generated reformatted video file and makes it accessible to users via a download link, which is communicated to users via email and / or their dashboard.
[0536] Step 11:
[0537] The user can then download the converted video file to their device using the provided download link, then burn it to a DVD or play it on their smartphone.
[0538] Step 12:
[0539] The device provides a communication application for sharing downloaded video files with family and friends. The user can use this application to upload video files to a cloud service or send them directly to family and friends.
[0540] Example 2
[0541] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0542] Due to aging and differences in formats, it is difficult to play the video and audio from old videotapes on modern devices. Furthermore, these videos suffer from noise and low resolution, resulting in a poor viewing experience. Furthermore, there is a lack of technology to recognize user emotions and enrich the visual expression. Therefore, current technology is required to improve the image and sound quality of old videotapes, convert them into formats playable on modern devices, and even add effects that respond to user emotions.
[0543] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for receiving video data digitized from an old video tape, a means for improving the image quality of the digitized video data using a generative AI model, and a means for improving the sound quality of the audio of the digitized video data using a generative AI model. This makes it possible to analyze the video data with high image quality and sound quality and recognize the user's emotions. In addition, by including a means for adding filters and effects to the video data based on the recognized emotions, the visual expression can be enriched.
[0544] "Digitization" is the process of converting analog data into digital data.
[0545] A "generative AI model" is an algorithm that uses artificial intelligence to generate or transform data for a specific task using techniques such as deep learning.
[0546] "High-definition" refers to the process of improving the resolution and image quality of video data.
[0547] "High quality audio" is the process of improving the quality of audio data and removing noise.
[0548] "Emotion recognition" is a technology that analyzes facial expressions and vocal tones of people contained in video data to detect specific emotional states (such as joy, anger, sadness, or happiness).
[0549] A "filter" is a function for applying a specific effect to video data.
[0550] "Effects" is a function that adds special visual effects to video data.
[0551] "Reformatting" is the process of converting video data into a different form or format.
[0552] "Providing" refers to making the generated data available to users.
[0553] "Server" means a central control unit for receiving, storing, processing, and providing video data.
[0554] "User" means the person or end user who utilizes the system to upload video data and download the final product.
[0555] This invention is a system that uses a generative AI model to convert video data digitized from old videotapes into high-quality image and sound, combines it with an emotion engine that recognizes the user's emotions, processes and reformats the data according to the user's emotions, and finally dubs it into a format that can be played on DVDs, smartphones, etc. Specific steps for implementing this invention are described below.
[0556] Video data digitization and uploading
[0557] A user converts old videotapes (e.g., VHS or miniDV) into digital files using dedicated digitizing equipment. This can be done using a home video capture device. The digitized video data is saved on the device. The user then opens a specific website or application on the device and uploads the saved video data to a server. The website provides an interface for selecting and uploading files, and the user selects and uploads the digital data.
[0558] High-quality image and sound processing
[0559] The server receives and temporarily stores video files uploaded by users. The server then launches a generative AI model to process the uploaded video data. This generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs similar noise reduction and sound quality clearing on the audio track to improve sound quality. For example, the generative AI model upscales a 480p home video to 1080p and generates a noise-free video. It also performs processing to remove background noise from the audio track and improve sound quality.
[0560] User Emotion Recognition
[0561] The server analyzes the high-quality video data and invokes an emotion engine to recognize the user's emotions. The emotion engine analyzes the facial expressions and tone of voice of the people in the video data to detect emotions such as smiling or crying. Based on these emotions, filters and effects are added to specific parts of the video. For example, if the emotion engine analyzes a scene in a home video where a child blows out a birthday cake and detects a happy expression, it adds a special effect to that scene.
[0562] Reformatting and providing data
[0563] The server converts the processed video file with high image and sound quality into the format specified by the user (e.g., DVD, MP4, etc.). The converted data is then saved again, and a download link is generated and sent to the user. The user then downloads the video data to their device via the link. The user can then burn the downloaded file to a DVD or play it on their smartphone. They can also share the video with family and friends using cloud services or communication applications. For example, a user can download an ISO file provided by the server and burn it to a DVD using their home DVD burning software. They can then play the DVD on a home DVD player and enjoy it with their family. They can also upload the converted MP4 file to cloud storage and share it with family and friends using a specific communication application.
[0564] Prompt Sentence Examples
[0565] 1. "What are the specific steps to digitize and upload my old VHS tapes?"
[0566] 2. "How can I use a generative AI model to convert uploaded videos into high-quality video and audio?"
[0567] 3. "Please explain how you can analyze video data, recognize user emotions, and add specific effects."
[0568] The above is a specific embodiment for implementing the present invention. The present invention is a system for converting old videotape digital data into high-quality data, processing it according to the user's emotions, and making it easy to play and share on modern devices.
[0569] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0570] Step 1: Digitize and upload your video data
[0571] A user converts old videotapes (e.g., VHS or miniDV) into digital files using dedicated digitizing equipment (e.g., a home video capture device). The converted digital video data is stored on the device. Next, the user opens a specific website or application on the device to upload the saved digital data to a server. Through the file upload interface, the user selects the video data and clicks the upload button. The server receives this video data and temporarily stores it. The input is a digitized video file, and the output is video data temporarily stored on the server.
[0572] Step 2: High-quality image and sound processing
[0573] The server retrieves the stored video data and launches a generative AI model. When the video data is input into the generative AI model, it uses deep learning technology to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs noise reduction and sound quality clearing on the audio track. As a specific example, the generative AI model upscales a 480p home video to 1080p and generates a noise-free video. At the same time, it removes background noise from the audio track and performs processing to clear the sound quality. The input is temporarily stored video data, and the output is video data with high image quality and sound quality.
[0574] Step 3: Recognizing user emotions
[0575] The server receives high-quality video data from the generative AI model and passes it to the emotion engine. When video data is input into the emotion engine, it analyzes the facial expressions and voice tones of the people in the video to detect emotions such as smiling or crying. Based on the emotions recognized by the emotion engine, filters and effects are added to specific scenes. For example, if the emotion engine analyzes a scene of a child blowing out a birthday cake and detects a joyful expression, it adds a special effect to that scene. The input is high-quality video data, and the output is video data with emotion recognition and effects applied.
[0576] Step 4: Reformat and provide data
[0577] Once processing is complete, the server converts the video data, which has undergone high-quality image and sound quality enhancement and emotion recognition processing, into the format specified by the user (e.g., DVD, MP4, etc.). The converted data is then saved back to the server, and a download link is generated. The user then downloads the converted video data to their device via this download link. For example, a user downloads an ISO file provided by the server and burns it to a DVD using their home DVD burning software. The user then plays the DVD on a home DVD player and watches it with their family. Alternatively, the user can upload the converted MP4 file to cloud storage and share it with family and friends using a specific communication application. The input is video data with emotion recognition and effects applied, and the output is video data converted to the user-specified format.
[0578] The above describes the specific operations and data flow for each processing step. This system allows us to convert old videotapes to high quality, add visual effects based on the user's emotions, and provide them in a format that can be played and shared on modern devices.
[0579] (Application example 2)
[0580] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0581] Video and audio stored on old videotapes are not suitable for playback on modern devices with high image and sound quality. Furthermore, digitized data from these videotapes suffers from degradation of the original image and sound quality, making it difficult to achieve visual and auditory satisfaction. Furthermore, the video recorded on home videotapes is filled with people's memories and emotions, and there is a demand for content that reflects those emotions. To solve these issues, not only is it necessary to improve the quality of video data, but it is also necessary to edit it in a way that reflects the user's emotions.
[0582] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0583] In this invention, the server includes means for receiving video data digitized from old videotapes, means for improving the image quality of the digitized video data using a generative AI model, means for improving the audio quality of the digitized video data using a generative AI model, means for reformatting the video data with improved image quality and audio quality, means for providing the reformatted video data to a user, means for recognizing the emotions of people in the video data using an emotion engine and adding effects according to the recognized emotions, and means for outputting the video data in a specified format. This makes it possible to convert data from old videotapes into high-quality image and audio and provide it as video that reflects the user's emotions.
[0584] "Old videotapes" are magnetic tape media that record video and audio in a non-digital analog format, such as VHS or miniDV.
[0585] "Digitalized video data" means video and audio data recorded on analog videotape that has been converted into a digital format that can be processed and stored on an electronic device.
[0586] A "generative AI model" is an artificial intelligence model designed based on deep learning technology, and is an algorithm that analyzes and processes input data to generate specific results.
[0587] "High-definition" refers to the process of processing video data that has been degraded by low resolution or noise using a generative AI model to convert it into high-resolution, clear video.
[0588] "High-quality sound" refers to processing audio data containing background noise and degraded sound using a generative AI model to convert it into clear, high-quality sound.
[0589] "Reformatting" refers to the process of converting digitized video data into a specific file format or designated media format.
[0590] An "emotion engine" is an artificial intelligence technology that analyzes facial expressions and tone of voice of people contained in video and audio data and recognizes their emotions.
[0591] "Effects" refers to the process of adding specific visual or auditory effects to video or audio.
[0592] "Video data metadata" refers to information related to video data, such as supplementary information such as the date and time of filming, the location of filming, and the characters appearing in the video.
[0593] The "specified format" refers to the particular file or media format selected by the user for the final output of the video data.
[0594] This invention is a system that uses generative AI to convert video data digitized from old videotapes into high-quality image and sound, then combines it with an emotion engine that recognizes the user's emotions to add specific effects to the video data, and finally provides it in a format specified by the user.
[0595] Overall system configuration
[0596] The system mainly includes the following components:
[0597] 1. Digitization tools: Video capture devices to generate digital data from old videotapes.
[0598] 2. Data receiving means: A terminal (e.g., a smartphone or PC) that provides an interface for uploading digitized video data to a server.
[0599] 3. Generative AI model: A deep learning model for processing digital video and audio data to produce high-quality images and sounds.
[0600] 4. Emotion Engine: An artificial intelligence engine to analyze the emotions of people in video data and add specific effects.
[0601] 5. Reformatting means: A device that converts video data with high image quality, high sound quality, and added emotional effects into a format specified by the user (e.g., MP4, DVD ISO).
[0602] 6. Data providing means: A server that generates a download link to provide the final video data to the user.
[0603] What the program does
[0604] 1. Upload video data:
[0605] The user starts a dedicated application on a device (such as a smartphone or PC) and uploads the digitized files of old videotapes to the server. At this stage, a file selection interface is provided.
[0606] 2. Processing by generative AI models:
[0607] The server temporarily stores the uploaded video data and uses a generative AI model to improve the video and audio quality. This model is based on deep learning technology using TensorFlow. Specifically, it improves the video resolution, removes noise and blur, and performs color correction. It also adjusts the audio track to remove background noise and ensure clear sound quality.
[0608] 3. Emotion engine processing:
[0609] The video data, which has been enhanced in quality by the generative AI model, is then analyzed using an emotion engine. The emotion engine uses OpenCV and other technologies to analyze the facial expressions and tone of voice of the people in the video, and adds specific effects depending on the emotions it recognizes. For example, in a video of a birthday party, a congratulatory effect can be added to scenes where a child is happy.
[0610] 4. Data Reformatting and Provision:
[0611] The final processed video data is converted into the format specified by the user (e.g. MP4 or DVD ISO). The server saves the converted video data in storage and notifies the user of a download link. Using this link, the user can download the video data and save it on their device.
[0612] Prompt Sentence Examples
[0613] "Convert your 480p home videos to high-quality 1080p footage, remove background noise and improve the clarity of your audio tracks."
[0614] As described above, the present invention is a system that converts digital data from old videotapes into high-quality data and adds effects that correspond to the user's emotions, thereby providing memorable videos in a more attractive form.
[0615] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0616] Step 1:
[0617] Users launch a dedicated application on their device (smartphone or PC), select the digitized files of old videotapes, and upload them to the server.
[0618] Input: Digitized video data file
[0619] Process: The user selects a file using the file selection interface and presses the "Upload" button.
[0620] Output: Video data is uploaded to the server and temporarily stored.
[0621] Step 2:
[0622] The server temporarily stores the uploaded video data and uses generative AI models to improve the video and audio quality.
[0623] Input: Uploaded video data
[0624] Processing: Generative AI models are used to improve video resolution, remove noise and blur, and perform color correction. Specifically, deep learning models using TensorFlow perform frame-by-frame processing. Audio tracks are also adjusted to remove background noise and ensure clarity.
[0625] Output: High-quality video data
[0626] Step 3:
[0627] The emotion engine analyzes high-quality video data and audio to recognize people's emotions.
[0628] Input: High-quality video data
[0629] Processing: An emotion engine (using OpenCV, etc.) is used to analyze facial expressions and vocal tones of people in the video to recognize specific emotions.
[0630] Output: Metadata about the recognized emotion
[0631] Step 4:
[0632] The server adds specific effects to the video data according to the emotions recognized by the emotion engine.
[0633] Input: High-quality video data and emotion recognition metadata
[0634] Processing: Add effects to specific scenes based on the recognized emotion. For example, add a blessing effect to a scene where the emotion is recognized as "joy."
[0635] Output: Video data with effects added
[0636] Step 5:
[0637] The server then converts the final processed video data into the format specified by the user and generates a download link.
[0638] Input: Video data with effects added
[0639] Processing: Convert the data to the desired format (e.g. MP4, DVD ISO) and save it to storage, using a reformatting tool to convert it to the appropriate file format, and generate a link for the user to download it.
[0640] Output: Video data in the specified format and a download link
[0641] Step 6:
[0642] The user uses the provided download link to download the final video data to their terminal.
[0643] Input: Download link
[0644] Processing: The user clicks on the notified link and downloads the video data.
[0645] Output: The final video data saved on your device
[0646] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0647] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0648] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0649] [Third embodiment]
[0650] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0651] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0652] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0653] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0654] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0655] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0656] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0657] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0658] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0659] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0660] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0661] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0662] The present invention is a system for converting video data digitized from old videotapes into high-quality image and sound using generation AI, and dubbing it into a format that can be played on DVDs, smartphones, etc. The main modes for implementing the present invention are described in detail below.
[0663] Uploading video data
[0664] Users convert old videotapes (e.g., VHS or miniDV) into digital files using a dedicated digitizing device and save them on their device. Next, they open a specific website or application on their device to upload the saved digital data to a server. The web page provides an interface for selecting and uploading files.
[0665] Example: A user uses a home video capture device to digitize their VHS tapes and save them on their computer. Then, they access a dedicated website, select the saved video files, and upload them.
[0666] High-quality image and sound processing
[0667] The server receives and temporarily stores the uploaded video file. It then processes the received data using a generative AI model. The generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs noise reduction and clearing on the audio track to improve sound quality.
[0668] Example: A generative AI model upscales an original 480p home video to 1080p and produces a noise-reduced video. The audio track is processed to remove background noise and improve audio clarity.
[0669] Reformatting and providing data
[0670] Once the processing is complete, the server converts the high-quality video file to the format of the user's choice (e.g. DVD, MP4, etc.). The converted data is then re-saved and a download link is generated. This link is then communicated to the user for easy access.
[0671] Users can download the converted video data using the provided link, burn it to a disc for playback on a home DVD player, play it on a smartphone, or share it with family and friends using a communication application or cloud service for data sharing.
[0672] For example, a user downloads an ISO file provided by a server and burns it to a DVD disc using home DVD burning software. Then, the user plays the DVD on a home DVD player and watches it with his or her family. Alternatively, the user uploads the converted MP4 file to cloud storage and shares it with family and friends using a specific communication application.
[0673] Thus, the present invention is a system that uses generative AI to convert the video and audio from old videotapes into high-quality footage that users can easily play and share on modern devices.
[0674] The processing flow will be explained below.
[0675] Step 1:
[0676] Users convert old videotapes (VHS, miniDV, etc.) into digital data using dedicated digitizing equipment. The digitized video data is then stored on the user's device.
[0677] Step 2:
[0678] The terminal provides an interface for uploading the digitized video data to the server through a dedicated website or application. The user selects a file using the terminal interface and clicks the "upload" button.
[0679] Step 3:
[0680] The server receives the video data uploaded by the user and temporarily stores it. When storing it, it analyzes and extracts the metadata of the video data (resolution, frame rate, audio format, etc.).
[0681] Step 4:
[0682] The server then calls the generative AI model, which uses deep learning techniques to enhance the image quality of each frame, remove noise and blur, and perform color correction.
[0683] Step 5:
[0684] The server uses a generative AI model to process the extracted audio track to improve its quality, specifically by removing noise, canceling echoes, and adjusting pitch and clearing to improve sound quality.
[0685] Step 6:
[0686] The server recombines the high-quality video and audio data to generate a new video file, which is then reformatted to the user's desired format (e.g., DVD or MP4).
[0687] Step 7:
[0688] The server stores the generated reformatted video file and makes it accessible to users via a download link, which is communicated to users via email and / or their dashboard.
[0689] Step 8:
[0690] The user can then download the converted video file to their device using the provided download link, then burn it to a DVD or play it on their smartphone.
[0691] Step 9:
[0692] The device provides a communication application for sharing downloaded video files with family and friends. The user can use this application to upload video files to a cloud service or send them directly to family and friends.
[0693] Example 1
[0694] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0695] Many households currently hold old analog recording media (such as VHS tapes and miniDV) that have been recorded in the past, but they face the problem of being difficult to play back on modern devices. Another issue is the deterioration of these analog recording media, which leads to a decline in the quality of the video and audio. There is a need to solve these problems and convert valuable video assets from the past into high-quality digital data that can be easily played back and shared on modern devices.
[0696] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0697] In this invention, the server includes means for receiving digitized video data from old analog recording media, means for converting the data into digital files using dedicated digitizing equipment, means for improving the image quality of the digitized video data using a generative AI model, means for improving the audio quality of the digitized video data using a generative AI model, means for converting the improved image and audio quality video data into a format specified by a user, and means for providing the reformatted video data to a user. This allows the video and audio from old analog recording media to be converted into high-quality digital data that can be easily played and shared on modern devices.
[0698] An "analog recording medium" is a medium that records video and audio as analog signals, such as VHS tapes and miniDV.
[0699] "Digitalization" means converting video and audio recorded as analog signals into digital signals.
[0700] A "generative AI model" is an artificial intelligence model that uses machine learning and deep learning and is used to improve the quality of video and audio.
[0701] "High image quality" refers to processing to improve the image quality of video, and includes techniques such as improving resolution, removing noise, and correcting colors.
[0702] "High-quality sound" refers to processing to improve the quality of audio, and includes techniques such as noise removal and acoustic clearing.
[0703] "Formatting" refers to converting digital data into a format that can be played on a specific playback device or software, examples of which include MP4 and DVD ISO files.
[0704] A "download link" is a link that allows a user to obtain digital data via the Internet.
[0705] A "digital file" is a data file recorded as a digital signal, examples of which include video files and audio files.
[0706] "Specialized digitizing equipment" means equipment designed to digitize analog recording media, examples of which include video capture devices.
[0707] "Reformatting" means converting processed digital data into a different format.
[0708] The present invention is a system for digitizing video data from old analog recording media, converting it into high-quality image and sound using a generative AI model, and dubbing it into a format that can be played and shared on modern devices. The main aspects of implementing the present invention are described in detail below.
[0709] Video data digitization and uploading
[0710] Users use specialized digitizing equipment (e.g., video capture devices) to convert old analog recording media (e.g., VHS tapes, miniDV) into digital files. They save these digital files to their own devices and then access a specific website or application to upload the saved digital files to a server. An upload interface is provided, which users use to select files and start uploading.
[0711] Example: A user uses a home video capture device to digitize their VHS tapes and save them on their computer. Then, they access a dedicated website, select the saved video files, and upload them.
[0712] High-quality image and sound processing using generative AI models
[0713] The server receives and temporarily stores the uploaded video file. It then processes the video data using a generative AI model (e.g., Super Resolution GAN, Noise2Noise). The generative AI model performs the following tasks:
[0714] 1. Increase the resolution of each frame (using Super Resolution technology).
[0715] 2. Remove noise and blur (using De-Noising technology).
[0716] 3. Perform color correction (using Color Stability technology).
[0717] 4. The audio track is also subjected to noise removal and clearing to improve sound quality (using Audio Enhancement technology).
[0718] Example: A generative AI model upscales an original 480p home video to 1080p and produces a noise-reduced video. The audio track is processed to remove background noise and improve audio clarity.
[0719] Reformatting and providing data
[0720] Once processing is complete, the server uses software such as ffmpeg or HandBrake to convert the high-quality video file into the format of the user's choice (e.g., DVD ISO file, MP4 format). The converted data is then stored back on the server and a download link is generated. This link is then posted to the user for easy access.
[0721] User download and sharing of data
[0722] Users can download the converted video data using the provided link. They can then burn the downloaded data to a disc using home DVD burning software (e.g., Nero Burning ROM) or play it on their smartphone. They can also share the data with family and friends using cloud services or communication applications.
[0723] For example, a user downloads an ISO file provided by a server and burns it to a DVD disc using home DVD burning software. Then, the user plays the DVD on a home DVD player and watches it with his or her family. Alternatively, the user uploads the converted MP4 file to cloud storage and shares it with family and friends using a specific communication application.
[0724] Prompt Sentence Examples
[0725] "I have digitized some old VHS tapes. I would like to convert these files to high-quality MP4 format that can be played on my smartphone. Please use your generative AI model to convert them."
[0726] Thus, the present invention is a system that uses generative AI to convert video and audio from old analog recording media into high-quality content that users can easily play and share on modern devices.
[0727] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0728] Step 1: User digitizes video data
[0729] Users convert old analog recording media (e.g., VHS tapes, miniDV) into digital files using dedicated digitizing equipment (e.g., video capture devices). They start software on a device connected to the digitizing equipment, press the play button, and the analog video and audio are converted into digital signals, which are then saved on the device as digital files (e.g., .mp4, .avi).
[0730] Input: Analog recording media
[0731] Output: Digitized video data file
[0732] What it does: A user inserts a VHS tape into a video capture device, runs the software, and begins recording. Once the recording is complete, it is saved as a digital file on the device.
[0733] Step 2: User uploads digital data
[0734] The user accesses a specific website or application on his / her own terminal to upload the digital data saved in the previous step to the server, selects the file from the upload interface of the website or application, and clicks the upload button to start the file transfer.
[0735] Input: Digitized video data file
[0736] Output: Video data file saved on the server
[0737] Specific operation: The user opens a browser, accesses the specified URL, logs in, selects the digitized video file in the file upload interface, and clicks the upload button.
[0738] Step 3: High-quality image and sound processing by the server
[0739] The server receives the uploaded video file, temporarily stores it, and then uses a generative AI model (e.g., Super Resolution GAN, Noise2Noise) to enhance the image and sound quality of the video data.
[0740] 1. Apply Super Resolution technology to improve the resolution of each frame.
[0741] 2. De-Noising technology is applied to remove noise and blur.
[0742] 3. Color Stability technology is applied to perform color correction.
[0743] 4. Apply Audio Enhancement technology to denoise and clear the audio track.
[0744] Input: Uploaded video file
[0745] Output: High-quality video and audio files
[0746] What happens: The server inputs a prompt into the generative AI model, requesting, "Please improve the image and sound quality of this video." The generative AI model then begins processing, optimizing each frame and audio track.
[0747] Step 4: Reformatting Data by the Server
[0748] The processed video file is converted to the format specified by the user. The server uses tools such as ffmpeg or HandBrake to reformat the file and saves the newly generated file on the server.
[0749] Input: High-quality video files
[0750] Output: Video file in the specified format
[0751] Specific operation: The server runs the ffmpeg command to convert the video file to MP4 format or a DVD ISO file, and then saves the converted file.
[0752] Step 5: Generate and notify the download link
[0753] The server generates a download link for the reformatted video data and notifies the user via email or web notification.
[0754] Input: Reformatted video file
[0755] Output: Download link
[0756] What happens: The server generates a download link and sends it to the user's email address, or displays a notification on the web interface.
[0757] Step 6: Users download and share their data
[0758] The user can then use the provided link to download the converted video data, which can then be burned to a disc using home DVD burning software, played on a smartphone, or shared via cloud services or communication applications.
[0759] Input: Download link
[0760] Output: High-quality video data stored on home devices or cloud storage
[0761] What it does: A user clicks the download link and saves the file to their device. They then use DVD burning software to burn the ISO file to a disc and play it on a DVD player. Alternatively, they can play the downloaded MP4 file on their smartphone. They can also upload it to cloud storage and share it with family and friends.
[0762] (Application example 1)
[0763] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0764] One challenge is the difficulty of playing back memorable footage stored on old videotapes with high image and sound quality on modern devices. Additionally, there are limited ways for users to easily digitize these videos and obtain them in the format of their choice. Furthermore, traditional methods for remastering videos require specialized knowledge and software, making them inaccessible to the average user.
[0765] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0766] In this invention, the server includes means for receiving video data digitized from old videotapes, means for improving the image quality of the digitized video data using a generative AI model, means for improving the audio quality of the digitized video data using a generative AI model, means for reformatting the video data with improved image quality and audio quality, means for providing the reformatted video data to a user, and means for improving the image quality and audio quality of the video data on the cloud server and making the video data converted into a format selected by the user available for download. This allows users to easily convert old video footage into high quality and obtain it in a format playable on modern devices.
[0767] "Old videotapes" are analog storage media such as VHS and miniDV, and are tape devices used to store video and audio data.
[0768] "Digital video data" refers to video and audio data that has been converted from analog videotape and stored as a digital data file on a computer or other device.
[0769] A "generative AI model" is a type of artificial intelligence based on deep learning technology that generates new data based on input data. It has the potential to improve the quality of video and audio.
[0770] "High-definition" refers to the process of improving the image quality by increasing the resolution of existing images, removing noise and blur, and performing color correction.
[0771] "High-quality sound enhancement" is an acoustic process that removes background noise from the audio track and improves intelligibility and clarity of the sound.
[0772] "Reformatting" refers to converting the processed data into a different file format to optimize it for a specific playback device or purpose, such as MP4, MKV, or DVD.
[0773] A "cloud server" is a remote server available via the Internet, a computer system that provides large-volume data processing and storage.
[0774] A "format" refers to the data storage format or structure, and is a rule for storing digital data in a specific way so that it can be played or used.
[0775] This invention is a system that uses a generative AI model to convert video data digitized from old videotapes into high-quality image and sound, and dubs it into a format playable on modern devices. The system is designed to allow users to easily upload video data and download high-quality video data processed by the generative AI model.
[0776] Uploading video data
[0777] First, a user digitizes old videotapes (e.g., VHS or miniDV) using a home video capture device and saves them as digital files on their device. Then, the user opens a dedicated website or application and uploads the digitized video files to a server. The user selects the files using an intuitive interface and starts the upload. For example, a user can digitize a VHS tape at home, save it on their computer, and then access a dedicated website to select and upload the video files.
[0778] High-quality image and sound processing
[0779] The server receives and temporarily stores the uploaded video file. It then processes the video data using a generative AI model. The generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It can also denoise and clear the audio track, resulting in a higher quality sound. For example, it can upscale an original 480p home video to 1080p and remove background noise from the audio track for a clearer sound.
[0780] Reformatting and providing data
[0781] After the generative AI model completes the high-quality image and sound processing, the server reformats the processed video data. The user can then select the desired output format (e.g., DVD or MP4) and download the converted video data. The user obtains the video data through the generated download link. The user downloads the provided ISO file, burns it to a DVD using home DVD burning software, and plays it on a home DVD player. The converted MP4 file can also be uploaded to cloud storage and shared with family and friends through certain communication applications. An example prompt is, "Upload your old VHS home video and upscale it to high-quality 1080p with clear audio. This video will be available for download in MP4 format."
[0782] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0783] Step 1:
[0784] A user digitizes old videotapes.
[0785] Input: Video tape, video capture device
[0786] Output: Digital file
[0787] What it does: A user uses a home video capture device to convert video and audio data from an old videotape into digital files and save them on their device.
[0788] Step 2:
[0789] A user uploads a digitized video file to a server.
[0790] Input: Digital file, website upload interface
[0791] Output: Video file temporarily saved on the server
[0792] How it works: A user accesses a dedicated website or application, clicks the upload button, selects a digitized video file, and uploads it to the server.
[0793] Step 3:
[0794] The server processes the video data to improve its image and sound quality.
[0795] Input: Digital files, generative AI models
[0796] Output: High-quality video data
[0797] How it works: The server processes the uploaded video file using a generative AI model. Specifically, it improves the video resolution, removes noise and blur, and performs color correction. It also denoises and clears the audio track. The result is video and audio of higher quality than the original.
[0798] Step 4:
[0799] The server reformats the processed video data.
[0800] Input: High-quality video data, user-selectable format
[0801] Output: Video file in the specified format
[0802] What happens: The server converts the processed video data into the format selected by the user (e.g. MP4, DVD, etc.) using the appropriate encoding software.
[0803] Step 5:
[0804] The server generates a download link for the video data and provides it to the user.
[0805] Input: Format converted video file
[0806] Output: Download link
[0807] Specific operation: The server hosts the converted video file and generates a download link that users can access. The server notifies users of this link so that they can easily obtain the data.
[0808] Step 6:
[0809] Users can download and use high-quality video files.
[0810] Input: Download link
[0811] Output: Locally saved video file
[0812] Specific operation: Users can click the download link provided to download the high-quality video file to their device and play or share it.
[0813] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0814] The present invention is a system that uses generation AI to convert video data digitized from old videotapes into high-quality image and sound, combines it with an emotion engine that recognizes the user's emotions, processes and reformats the data according to the user's emotions, and finally dubs the data into a format that can be played on DVDs, smartphones, etc. The main modes for implementing the present invention are described in detail below.
[0815] Uploading video data
[0816] Users convert old videotapes (e.g., VHS or miniDV) into digital files using a dedicated digitizing device and save them on their device. Next, they open a specific website or application on their device to upload the saved digital data to a server. The web page provides an interface for selecting and uploading files.
[0817] Example: A user uses a home video capture device to digitize their VHS tapes and save them on their computer. Then, they access a dedicated website, select the saved video files, and upload them.
[0818] High-quality image and sound processing
[0819] The server receives and temporarily stores the uploaded video file. It then processes the received data using a generative AI model. The generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs noise reduction and clearing on the audio track to improve sound quality.
[0820] Example: A generative AI model upscales an original 480p home video to 1080p and produces a noise-reduced video. The audio track is processed to remove background noise and improve audio clarity.
[0821] User Emotion Recognition
[0822] The server analyzes the high-quality video data and invokes an emotion engine to recognize the user's emotions. The emotion engine analyzes the facial expressions and tone of voice of the people in the video data to detect emotions such as smiling or crying. Based on this emotion, filters and effects are added to specific parts of the video.
[0823] Example: An emotion engine analyzes a scene in a home video where a child blows out a birthday cake, and if it detects a happy expression, it adds a special effect to the scene.
[0824] Reformatting and providing data
[0825] Once the processing is complete, the server converts the high-quality video file to the format of the user's choice (e.g. DVD, MP4, etc.). The converted data is then re-saved and a download link is generated. This link is then communicated to the user for easy access.
[0826] The user can then use the provided link to download the converted video data to their device. They can then burn the downloaded file to a DVD, play it on their smartphone, or share the video with family and friends using a communication app or cloud service for data sharing.
[0827] For example, a user downloads an ISO file provided by a server and burns it to a DVD disc using home DVD burning software. Then, the user plays the DVD on a home DVD player and watches it with his or her family. Alternatively, the user uploads the converted MP4 file to cloud storage and shares it with family and friends using a specific communication application.
[0828] Thus, the present invention is a system that uses generative AI and an emotion engine to convert the video and audio from old videotapes into high-quality footage, apply specific processing based on the user's emotions, and make it easy to play and share on modern devices.
[0829] The processing flow will be explained below.
[0830] Step 1:
[0831] Users convert old videotapes (VHS, miniDV, etc.) into digital data using dedicated digitizing equipment. The digitized video data is then stored on the user's device.
[0832] Step 2:
[0833] The terminal provides an interface for uploading the digitized video data to the server through a dedicated website or application. The user selects a file using the terminal interface and clicks the "upload" button.
[0834] Step 3:
[0835] The server receives the video data uploaded by the user and temporarily stores it. When storing it, it analyzes and extracts the metadata of the video data (resolution, frame rate, audio format, etc.).
[0836] Step 4:
[0837] The server then calls the generative AI model, which uses deep learning techniques to enhance the image quality of each frame, remove noise and blur, and perform color correction.
[0838] Step 5:
[0839] The server uses a generative AI model to process the extracted audio track to improve its quality, specifically by removing noise, canceling echoes, and adjusting pitch and clearing to improve sound quality.
[0840] Step 6:
[0841] The server recombines the high-quality video and audio data to generate a new video file, which is then reformatted to the format specified by the user (e.g., DVD format, MP4 format).
[0842] Step 7:
[0843] The server invokes the emotion engine to analyze the facial expressions and voices of people in the video data, and the emotion engine detects specific emotions, such as smiling or crying.
[0844] Step 8:
[0845] The server processes the video by adding filters and effects to specific parts of the video based on the emotions detected by the emotion engine. For example, a special effect is applied to scenes where the emotion of joy is detected.
[0846] Step 9:
[0847] The server reformats the emotion-processed, high-quality video data into a format specified by the user, and generates and saves the final video file.
[0848] Step 10:
[0849] The server stores the generated reformatted video file and makes it accessible to users via a download link, which is communicated to users via email and / or their dashboard.
[0850] Step 11:
[0851] The user can then download the converted video file to their device using the provided download link, then burn it to a DVD or play it on their smartphone.
[0852] Step 12:
[0853] The device provides a communication application for sharing downloaded video files with family and friends. The user can use this application to upload video files to a cloud service or send them directly to family and friends.
[0854] Example 2
[0855] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0856] Due to aging and differences in formats, it is difficult to play the video and audio from old videotapes on modern devices. Furthermore, these videos suffer from noise and low resolution, resulting in a poor viewing experience. Furthermore, there is a lack of technology to recognize user emotions and enrich the visual expression. Therefore, current technology is required to improve the image and sound quality of old videotapes, convert them into formats playable on modern devices, and even add effects that respond to user emotions.
[0857] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for receiving video data digitized from an old video tape, a means for improving the image quality of the digitized video data using a generative AI model, and a means for improving the sound quality of the audio of the digitized video data using a generative AI model. This makes it possible to analyze the video data with high image quality and sound quality and recognize the user's emotions. In addition, by including a means for adding filters and effects to the video data based on the recognized emotions, the visual expression can be enriched.
[0858] "Digitization" is the process of converting analog data into digital data.
[0859] A "generative AI model" is an algorithm that uses artificial intelligence to generate or transform data for a specific task using techniques such as deep learning.
[0860] "High-definition" refers to the process of improving the resolution and image quality of video data.
[0861] "High quality audio" is the process of improving the quality of audio data and removing noise.
[0862] "Emotion recognition" is a technology that analyzes facial expressions and vocal tones of people contained in video data to detect specific emotional states (such as joy, anger, sadness, or happiness).
[0863] A "filter" is a function for applying a specific effect to video data.
[0864] "Effects" is a function that adds special visual effects to video data.
[0865] "Reformatting" is the process of converting video data into a different form or format.
[0866] "Providing" refers to making the generated data available to users.
[0867] "Server" means a central control unit for receiving, storing, processing, and providing video data.
[0868] "User" means the person or end user who utilizes the system to upload video data and download the final product.
[0869] This invention is a system that uses a generative AI model to convert video data digitized from old videotapes into high-quality image and sound, combines it with an emotion engine that recognizes the user's emotions, processes and reformats the data according to the user's emotions, and finally dubs it into a format that can be played on DVDs, smartphones, etc. Specific steps for implementing this invention are described below.
[0870] Video data digitization and uploading
[0871] A user converts old videotapes (e.g., VHS or miniDV) into digital files using dedicated digitizing equipment. This can be done using a home video capture device. The digitized video data is saved on the device. The user then opens a specific website or application on the device and uploads the saved video data to a server. The website provides an interface for selecting and uploading files, and the user selects and uploads the digital data.
[0872] High-quality image and sound processing
[0873] The server receives and temporarily stores video files uploaded by users. The server then launches a generative AI model to process the uploaded video data. This generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs similar noise reduction and sound quality clearing on the audio track to improve sound quality. For example, the generative AI model upscales a 480p home video to 1080p and generates a noise-free video. It also performs processing to remove background noise from the audio track and improve sound quality.
[0874] User Emotion Recognition
[0875] The server analyzes the high-quality video data and invokes an emotion engine to recognize the user's emotions. The emotion engine analyzes the facial expressions and tone of voice of the people in the video data to detect emotions such as smiling or crying. Based on these emotions, filters and effects are added to specific parts of the video. For example, if the emotion engine analyzes a scene in a home video where a child blows out a birthday cake and detects a happy expression, it adds a special effect to that scene.
[0876] Reformatting and providing data
[0877] The server converts the processed video file with high image and sound quality into the format specified by the user (e.g., DVD, MP4, etc.). The converted data is then saved again, and a download link is generated and sent to the user. The user then downloads the video data to their device via the link. The user can then burn the downloaded file to a DVD or play it on their smartphone. They can also share the video with family and friends using cloud services or communication applications. For example, a user can download an ISO file provided by the server and burn it to a DVD using their home DVD burning software. They can then play the DVD on a home DVD player and enjoy it with their family. They can also upload the converted MP4 file to cloud storage and share it with family and friends using a specific communication application.
[0878] Prompt Sentence Examples
[0879] 1. "What are the specific steps to digitize and upload my old VHS tapes?"
[0880] 2. "How can I use a generative AI model to convert uploaded videos into high-quality video and audio?"
[0881] 3. "Please explain how you can analyze video data, recognize user emotions, and add specific effects."
[0882] The above is a specific embodiment for implementing the present invention. The present invention is a system for converting old videotape digital data into high-quality data, processing it according to the user's emotions, and making it easy to play and share on modern devices.
[0883] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0884] Step 1: Digitize and upload your video data
[0885] A user converts old videotapes (e.g., VHS or miniDV) into digital files using dedicated digitizing equipment (e.g., a home video capture device). The converted digital video data is stored on the device. Next, the user opens a specific website or application on the device to upload the saved digital data to a server. Through the file upload interface, the user selects the video data and clicks the upload button. The server receives this video data and temporarily stores it. The input is a digitized video file, and the output is video data temporarily stored on the server.
[0886] Step 2: High-quality image and sound processing
[0887] The server retrieves the stored video data and launches a generative AI model. When the video data is input into the generative AI model, it uses deep learning technology to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs noise reduction and sound quality clearing on the audio track. As a specific example, the generative AI model upscales a 480p home video to 1080p and generates a noise-free video. At the same time, it removes background noise from the audio track and performs processing to clear the sound quality. The input is temporarily stored video data, and the output is video data with high image quality and sound quality.
[0888] Step 3: Recognizing user emotions
[0889] The server receives high-quality video data from the generative AI model and passes it to the emotion engine. When video data is input into the emotion engine, it analyzes the facial expressions and voice tones of the people in the video to detect emotions such as smiling or crying. Based on the emotions recognized by the emotion engine, filters and effects are added to specific scenes. For example, if the emotion engine analyzes a scene of a child blowing out a birthday cake and detects a joyful expression, it adds a special effect to that scene. The input is high-quality video data, and the output is video data with emotion recognition and effects applied.
[0890] Step 4: Reformat and provide data
[0891] Once processing is complete, the server converts the video data, which has undergone high-quality image and sound quality enhancement and emotion recognition processing, into the format specified by the user (e.g., DVD, MP4, etc.). The converted data is then saved back to the server, and a download link is generated. The user then downloads the converted video data to their device via this download link. For example, a user downloads an ISO file provided by the server and burns it to a DVD using their home DVD burning software. The user then plays the DVD on a home DVD player and watches it with their family. Alternatively, the user can upload the converted MP4 file to cloud storage and share it with family and friends using a specific communication application. The input is video data with emotion recognition and effects applied, and the output is video data converted to the user-specified format.
[0892] The above describes the specific operations and data flow for each processing step. This system allows us to convert old videotapes to high quality, add visual effects based on the user's emotions, and provide them in a format that can be played and shared on modern devices.
[0893] (Application example 2)
[0894] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0895] Video and audio stored on old videotapes are not suitable for playback on modern devices with high image and sound quality. Furthermore, digitized data from these videotapes suffers from degradation of the original image and sound quality, making it difficult to achieve visual and auditory satisfaction. Furthermore, the video recorded on home videotapes is filled with people's memories and emotions, and there is a demand for content that reflects those emotions. To solve these issues, not only is it necessary to improve the quality of video data, but it is also necessary to edit it in a way that reflects the user's emotions.
[0896] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0897] In this invention, the server includes means for receiving video data digitized from old videotapes, means for improving the image quality of the digitized video data using a generative AI model, means for improving the audio quality of the digitized video data using a generative AI model, means for reformatting the video data with improved image quality and audio quality, means for providing the reformatted video data to a user, means for recognizing the emotions of people in the video data using an emotion engine and adding effects according to the recognized emotions, and means for outputting the video data in a specified format. This makes it possible to convert data from old videotapes into high-quality image and audio and provide it as video that reflects the user's emotions.
[0898] "Old videotapes" are magnetic tape media that record video and audio in a non-digital analog format, such as VHS or miniDV.
[0899] "Digitalized video data" means video and audio data recorded on analog videotape that has been converted into a digital format that can be processed and stored on an electronic device.
[0900] A "generative AI model" is an artificial intelligence model designed based on deep learning technology, and is an algorithm that analyzes and processes input data to generate specific results.
[0901] "High-definition" refers to the process of processing video data that has been degraded by low resolution or noise using a generative AI model to convert it into high-resolution, clear video.
[0902] "High-quality sound" refers to processing audio data containing background noise and degraded sound using a generative AI model to convert it into clear, high-quality sound.
[0903] "Reformatting" refers to the process of converting digitized video data into a specific file format or designated media format.
[0904] An "emotion engine" is an artificial intelligence technology that analyzes facial expressions and tone of voice of people contained in video and audio data and recognizes their emotions.
[0905] "Effects" refers to the process of adding specific visual or auditory effects to video or audio.
[0906] "Video data metadata" refers to information related to video data, such as supplementary information such as the date and time of filming, the location of filming, and the characters appearing in the video.
[0907] The "specified format" refers to the particular file or media format selected by the user for the final output of the video data.
[0908] This invention is a system that uses generative AI to convert video data digitized from old videotapes into high-quality image and sound, then combines it with an emotion engine that recognizes the user's emotions to add specific effects to the video data, and finally provides it in a format specified by the user.
[0909] Overall system configuration
[0910] The system mainly includes the following components:
[0911] 1. Digitization tools: Video capture devices to generate digital data from old videotapes.
[0912] 2. Data receiving means: A terminal (e.g., a smartphone or PC) that provides an interface for uploading digitized video data to a server.
[0913] 3. Generative AI model: A deep learning model for processing digital video and audio data to produce high-quality images and sounds.
[0914] 4. Emotion Engine: An artificial intelligence engine to analyze the emotions of people in video data and add specific effects.
[0915] 5. Reformatting means: A device that converts video data with high image quality, high sound quality, and added emotional effects into a format specified by the user (e.g., MP4, DVD ISO).
[0916] 6. Data providing means: A server that generates a download link to provide the final video data to the user.
[0917] What the program does
[0918] 1. Upload video data:
[0919] The user starts a dedicated application on a device (such as a smartphone or PC) and uploads the digitized files of old videotapes to the server. At this stage, a file selection interface is provided.
[0920] 2. Processing by generative AI models:
[0921] The server temporarily stores the uploaded video data and uses a generative AI model to improve the video and audio quality. This model is based on deep learning technology using TensorFlow. Specifically, it improves the video resolution, removes noise and blur, and performs color correction. It also adjusts the audio track to remove background noise and ensure clear sound quality.
[0922] 3. Emotion engine processing:
[0923] The video data, which has been enhanced in quality by the generative AI model, is then analyzed using an emotion engine. The emotion engine uses OpenCV and other technologies to analyze the facial expressions and tone of voice of the people in the video, and adds specific effects depending on the emotions it recognizes. For example, in a video of a birthday party, a congratulatory effect can be added to scenes where a child is happy.
[0924] 4. Data Reformatting and Provision:
[0925] The final processed video data is converted into the format specified by the user (e.g. MP4 or DVD ISO). The server saves the converted video data in storage and notifies the user of a download link. Using this link, the user can download the video data and save it on their device.
[0926] Prompt Sentence Examples
[0927] "Convert your 480p home videos to high-quality 1080p footage, remove background noise and improve the clarity of your audio tracks."
[0928] As described above, the present invention is a system that converts digital data from old videotapes into high-quality data and adds effects that correspond to the user's emotions, thereby providing memorable videos in a more attractive form.
[0929] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0930] Step 1:
[0931] Users launch a dedicated application on their device (smartphone or PC), select the digitized files of old videotapes, and upload them to the server.
[0932] Input: Digitized video data file
[0933] Process: The user selects a file using the file selection interface and presses the "Upload" button.
[0934] Output: Video data is uploaded to the server and temporarily stored.
[0935] Step 2:
[0936] The server temporarily stores the uploaded video data and uses generative AI models to improve the video and audio quality.
[0937] Input: Uploaded video data
[0938] Processing: Generative AI models are used to improve video resolution, remove noise and blur, and perform color correction. Specifically, deep learning models using TensorFlow perform frame-by-frame processing. Audio tracks are also adjusted to remove background noise and ensure clarity.
[0939] Output: High-quality video data
[0940] Step 3:
[0941] The emotion engine analyzes high-quality video data and audio to recognize people's emotions.
[0942] Input: High-quality video data
[0943] Processing: An emotion engine (using OpenCV, etc.) is used to analyze facial expressions and vocal tones of people in the video to recognize specific emotions.
[0944] Output: Metadata about the recognized emotion
[0945] Step 4:
[0946] The server adds specific effects to the video data according to the emotions recognized by the emotion engine.
[0947] Input: High-quality video data and emotion recognition metadata
[0948] Processing: Add effects to specific scenes based on the recognized emotion. For example, add a blessing effect to a scene where the emotion is recognized as "joy."
[0949] Output: Video data with effects added
[0950] Step 5:
[0951] The server then converts the final processed video data into the format specified by the user and generates a download link.
[0952] Input: Video data with effects added
[0953] Processing: Convert the data to the desired format (e.g. MP4, DVD ISO) and save it to storage, using a reformatting tool to convert it to the appropriate file format, and generate a link for the user to download it.
[0954] Output: Video data in the specified format and a download link
[0955] Step 6:
[0956] The user uses the provided download link to download the final video data to their terminal.
[0957] Input: Download link
[0958] Processing: The user clicks on the notified link and downloads the video data.
[0959] Output: The final video data saved on your device
[0960] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0961] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0962] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0963] [Fourth embodiment]
[0964] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0965] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0966] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0967] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0968] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0969] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0970] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0971] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0972] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0973] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0974] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0975] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0976] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0977] The present invention is a system for converting video data digitized from old videotapes into high-quality image and sound using generation AI, and dubbing it into a format that can be played on DVDs, smartphones, etc. The main modes for implementing the present invention are described in detail below.
[0978] Uploading video data
[0979] Users convert old videotapes (e.g., VHS or miniDV) into digital files using a dedicated digitizing device and save them on their device. Next, they open a specific website or application on their device to upload the saved digital data to a server. The web page provides an interface for selecting and uploading files.
[0980] Example: A user uses a home video capture device to digitize their VHS tapes and save them on their computer. Then, they access a dedicated website, select the saved video files, and upload them.
[0981] High-quality image and sound processing
[0982] The server receives and temporarily stores the uploaded video file. It then processes the received data using a generative AI model. The generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs noise reduction and clearing on the audio track to improve sound quality.
[0983] Example: A generative AI model upscales an original 480p home video to 1080p and produces a noise-reduced video. The audio track is processed to remove background noise and improve audio clarity.
[0984] Reformatting and providing data
[0985] Once the processing is complete, the server converts the high-quality video file to the format of the user's choice (e.g. DVD, MP4, etc.). The converted data is then re-saved and a download link is generated. This link is then communicated to the user for easy access.
[0986] Users can download the converted video data using the provided link, burn it to a disc for playback on a home DVD player, play it on a smartphone, or share it with family and friends using a communication application or cloud service for data sharing.
[0987] For example, a user downloads an ISO file provided by a server and burns it to a DVD disc using home DVD burning software. Then, the user plays the DVD on a home DVD player and watches it with his or her family. Alternatively, the user uploads the converted MP4 file to cloud storage and shares it with family and friends using a specific communication application.
[0988] Thus, the present invention is a system that uses generative AI to convert the video and audio from old videotapes into high-quality footage that users can easily play and share on modern devices.
[0989] The processing flow will be explained below.
[0990] Step 1:
[0991] Users convert old videotapes (VHS, miniDV, etc.) into digital data using dedicated digitizing equipment. The digitized video data is then stored on the user's device.
[0992] Step 2:
[0993] The terminal provides an interface for uploading the digitized video data to the server through a dedicated website or application. The user selects a file using the terminal interface and clicks the "upload" button.
[0994] Step 3:
[0995] The server receives the video data uploaded by the user and temporarily stores it. When storing it, it analyzes and extracts the metadata of the video data (resolution, frame rate, audio format, etc.).
[0996] Step 4:
[0997] The server then calls the generative AI model, which uses deep learning techniques to enhance the image quality of each frame, remove noise and blur, and perform color correction.
[0998] Step 5:
[0999] The server uses a generative AI model to process the extracted audio track to improve its quality, specifically by removing noise, canceling echoes, and adjusting pitch and clearing to improve sound quality.
[1000] Step 6:
[1001] The server recombines the high-quality video and audio data to generate a new video file, which is then reformatted to the user's desired format (e.g., DVD or MP4).
[1002] Step 7:
[1003] The server stores the generated reformatted video file and makes it accessible to users via a download link, which is communicated to users via email and / or their dashboard.
[1004] Step 8:
[1005] The user can then download the converted video file to their device using the provided download link, then burn it to a DVD or play it on their smartphone.
[1006] Step 9:
[1007] The device provides a communication application for sharing downloaded video files with family and friends. The user can use this application to upload video files to a cloud service or send them directly to family and friends.
[1008] Example 1
[1009] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1010] Many households currently hold old analog recording media (such as VHS tapes and miniDV) that have been recorded in the past, but they face the problem of being difficult to play back on modern devices. Another issue is the deterioration of these analog recording media, which leads to a decline in the quality of the video and audio. There is a need to solve these problems and convert valuable video assets from the past into high-quality digital data that can be easily played back and shared on modern devices.
[1011] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1012] In this invention, the server includes means for receiving digitized video data from old analog recording media, means for converting the data into digital files using dedicated digitizing equipment, means for improving the image quality of the digitized video data using a generative AI model, means for improving the audio quality of the digitized video data using a generative AI model, means for converting the improved image and audio quality video data into a format specified by a user, and means for providing the reformatted video data to a user. This allows the video and audio from old analog recording media to be converted into high-quality digital data that can be easily played and shared on modern devices.
[1013] An "analog recording medium" is a medium that records video and audio as analog signals, such as VHS tapes and miniDV.
[1014] "Digitalization" means converting video and audio recorded as analog signals into digital signals.
[1015] A "generative AI model" is an artificial intelligence model that uses machine learning and deep learning and is used to improve the quality of video and audio.
[1016] "High image quality" refers to processing to improve the image quality of video, and includes techniques such as improving resolution, removing noise, and correcting colors.
[1017] "High-quality sound" refers to processing to improve the quality of audio, and includes techniques such as noise removal and acoustic clearing.
[1018] "Formatting" refers to converting digital data into a format that can be played on a specific playback device or software, examples of which include MP4 and DVD ISO files.
[1019] A "download link" is a link that allows a user to obtain digital data via the Internet.
[1020] A "digital file" is a data file recorded as a digital signal, examples of which include video files and audio files.
[1021] "Specialized digitizing equipment" means equipment designed to digitize analog recording media, examples of which include video capture devices.
[1022] "Reformatting" means converting processed digital data into a different format.
[1023] The present invention is a system for digitizing video data from old analog recording media, converting it into high-quality image and sound using a generative AI model, and dubbing it into a format that can be played and shared on modern devices. The main aspects of implementing the present invention are described in detail below.
[1024] Video data digitization and uploading
[1025] Users use specialized digitizing equipment (e.g., video capture devices) to convert old analog recording media (e.g., VHS tapes, miniDV) into digital files. They save these digital files to their own devices and then access a specific website or application to upload the saved digital files to a server. An upload interface is provided, which users use to select files and start uploading.
[1026] Example: A user uses a home video capture device to digitize their VHS tapes and save them on their computer. Then, they access a dedicated website, select the saved video files, and upload them.
[1027] High-quality image and sound processing using generative AI models
[1028] The server receives and temporarily stores the uploaded video file. It then processes the video data using a generative AI model (e.g., Super Resolution GAN, Noise2Noise). The generative AI model performs the following tasks:
[1029] 1. Increase the resolution of each frame (using Super Resolution technology).
[1030] 2. Remove noise and blur (using De-Noising technology).
[1031] 3. Perform color correction (using Color Stability technology).
[1032] 4. The audio track is also subjected to noise removal and clearing to improve sound quality (using Audio Enhancement technology).
[1033] Example: A generative AI model upscales an original 480p home video to 1080p and produces a noise-reduced video. The audio track is processed to remove background noise and improve audio clarity.
[1034] Reformatting and providing data
[1035] Once processing is complete, the server uses software such as ffmpeg or HandBrake to convert the high-quality video file into the format of the user's choice (e.g., DVD ISO file, MP4 format). The converted data is then stored back on the server and a download link is generated. This link is then posted to the user for easy access.
[1036] User download and sharing of data
[1037] Users can download the converted video data using the provided link. They can then burn the downloaded data to a disc using home DVD burning software (e.g., Nero Burning ROM) or play it on their smartphone. They can also share the data with family and friends using cloud services or communication applications.
[1038] For example, a user downloads an ISO file provided by a server and burns it to a DVD disc using home DVD burning software. Then, the user plays the DVD on a home DVD player and watches it with his or her family. Alternatively, the user uploads the converted MP4 file to cloud storage and shares it with family and friends using a specific communication application.
[1039] Prompt Sentence Examples
[1040] "I have digitized some old VHS tapes. I would like to convert these files to high-quality MP4 format that can be played on my smartphone. Please use your generative AI model to convert them."
[1041] Thus, the present invention is a system that uses generative AI to convert video and audio from old analog recording media into high-quality content that users can easily play and share on modern devices.
[1042] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1043] Step 1: User digitizes video data
[1044] Users convert old analog recording media (e.g., VHS tapes, miniDV) into digital files using dedicated digitizing equipment (e.g., video capture devices). They start software on a device connected to the digitizing equipment, press the play button, and the analog video and audio are converted into digital signals, which are then saved on the device as digital files (e.g., .mp4, .avi).
[1045] Input: Analog recording media
[1046] Output: Digitized video data file
[1047] What it does: A user inserts a VHS tape into a video capture device, runs the software, and begins recording. Once the recording is complete, it is saved as a digital file on the device.
[1048] Step 2: User uploads digital data
[1049] The user accesses a specific website or application on his / her own terminal to upload the digital data saved in the previous step to the server, selects the file from the upload interface of the website or application, and clicks the upload button to start the file transfer.
[1050] Input: Digitized video data file
[1051] Output: Video data file saved on the server
[1052] Specific operation: The user opens a browser, accesses the specified URL, logs in, selects the digitized video file in the file upload interface, and clicks the upload button.
[1053] Step 3: High-quality image and sound processing by the server
[1054] The server receives the uploaded video file, temporarily stores it, and then uses a generative AI model (e.g., Super Resolution GAN, Noise2Noise) to enhance the image and sound quality of the video data.
[1055] 1. Apply Super Resolution technology to improve the resolution of each frame.
[1056] 2. De-Noising technology is applied to remove noise and blur.
[1057] 3. Color Stability technology is applied to perform color correction.
[1058] 4. Apply Audio Enhancement technology to denoise and clear the audio track.
[1059] Input: Uploaded video file
[1060] Output: High-quality video and audio files
[1061] What happens: The server inputs a prompt into the generative AI model, requesting, "Please improve the image and sound quality of this video." The generative AI model then begins processing, optimizing each frame and audio track.
[1062] Step 4: Reformatting Data by the Server
[1063] The processed video file is converted to the format specified by the user. The server uses tools such as ffmpeg or HandBrake to reformat the file and saves the newly generated file on the server.
[1064] Input: High-quality video files
[1065] Output: Video file in the specified format
[1066] Specific operation: The server runs the ffmpeg command to convert the video file to MP4 format or a DVD ISO file, and then saves the converted file.
[1067] Step 5: Generate and notify the download link
[1068] The server generates a download link for the reformatted video data and notifies the user via email or web notification.
[1069] Input: Reformatted video file
[1070] Output: Download link
[1071] What happens: The server generates a download link and sends it to the user's email address, or displays a notification on the web interface.
[1072] Step 6: Users download and share their data
[1073] The user can then use the provided link to download the converted video data, which can then be burned to a disc using home DVD burning software, played on a smartphone, or shared via cloud services or communication applications.
[1074] Input: Download link
[1075] Output: High-quality video data stored on home devices or cloud storage
[1076] What it does: A user clicks the download link and saves the file to their device. They then use DVD burning software to burn the ISO file to a disc and play it on a DVD player. Alternatively, they can play the downloaded MP4 file on their smartphone. They can also upload it to cloud storage and share it with family and friends.
[1077] (Application example 1)
[1078] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1079] One challenge is the difficulty of playing back memorable footage stored on old videotapes with high image and sound quality on modern devices. Additionally, there are limited ways for users to easily digitize these videos and obtain them in the format of their choice. Furthermore, traditional methods for remastering videos require specialized knowledge and software, making them inaccessible to the average user.
[1080] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1081] In this invention, the server includes means for receiving video data digitized from old videotapes, means for improving the image quality of the digitized video data using a generative AI model, means for improving the audio quality of the digitized video data using a generative AI model, means for reformatting the video data with improved image quality and audio quality, means for providing the reformatted video data to a user, and means for improving the image quality and audio quality of the video data on the cloud server and making the video data converted into a format selected by the user available for download. This allows users to easily convert old video footage into high quality and obtain it in a format playable on modern devices.
[1082] "Old videotapes" are analog storage media such as VHS and miniDV, and are tape devices used to store video and audio data.
[1083] "Digital video data" refers to video and audio data that has been converted from analog videotape and stored as a digital data file on a computer or other device.
[1084] A "generative AI model" is a type of artificial intelligence based on deep learning technology that generates new data based on input data. It has the potential to improve the quality of video and audio.
[1085] "High-definition" refers to the process of improving the image quality by increasing the resolution of existing images, removing noise and blur, and performing color correction.
[1086] "High-quality sound enhancement" is an acoustic process that removes background noise from the audio track and improves intelligibility and clarity of the sound.
[1087] "Reformatting" refers to converting the processed data into a different file format to optimize it for a specific playback device or purpose, such as MP4, MKV, or DVD.
[1088] A "cloud server" is a remote server available via the Internet, a computer system that provides large-volume data processing and storage.
[1089] A "format" refers to the data storage format or structure, and is a rule for storing digital data in a specific way so that it can be played or used.
[1090] This invention is a system that uses a generative AI model to convert video data digitized from old videotapes into high-quality image and sound, and dubs it into a format playable on modern devices. The system is designed to allow users to easily upload video data and download high-quality video data processed by the generative AI model.
[1091] Uploading video data
[1092] First, a user digitizes old videotapes (e.g., VHS or miniDV) using a home video capture device and saves them as digital files on their device. Then, the user opens a dedicated website or application and uploads the digitized video files to a server. The user selects the files using an intuitive interface and starts the upload. For example, a user can digitize a VHS tape at home, save it on their computer, and then access a dedicated website to select and upload the video files.
[1093] High-quality image and sound processing
[1094] The server receives and temporarily stores the uploaded video file. It then processes the video data using a generative AI model. The generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It can also denoise and clear the audio track, resulting in a higher quality sound. For example, it can upscale an original 480p home video to 1080p and remove background noise from the audio track for a clearer sound.
[1095] Reformatting and providing data
[1096] After the generative AI model completes the high-quality image and sound processing, the server reformats the processed video data. The user can then select the desired output format (e.g., DVD or MP4) and download the converted video data. The user obtains the video data through the generated download link. The user downloads the provided ISO file, burns it to a DVD using home DVD burning software, and plays it on a home DVD player. The converted MP4 file can also be uploaded to cloud storage and shared with family and friends through certain communication applications. An example prompt is, "Upload your old VHS home video and upscale it to high-quality 1080p with clear audio. This video will be available for download in MP4 format."
[1097] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1098] Step 1:
[1099] A user digitizes old videotapes.
[1100] Input: Video tape, video capture device
[1101] Output: Digital file
[1102] What it does: A user uses a home video capture device to convert video and audio data from an old videotape into digital files and save them on their device.
[1103] Step 2:
[1104] A user uploads a digitized video file to a server.
[1105] Input: Digital file, website upload interface
[1106] Output: Video file temporarily saved on the server
[1107] How it works: A user accesses a dedicated website or application, clicks the upload button, selects a digitized video file, and uploads it to the server.
[1108] Step 3:
[1109] The server processes the video data to improve its image and sound quality.
[1110] Input: Digital files, generative AI models
[1111] Output: High-quality video data
[1112] How it works: The server processes the uploaded video file using a generative AI model. Specifically, it improves the video resolution, removes noise and blur, and performs color correction. It also denoises and clears the audio track. The result is video and audio of higher quality than the original.
[1113] Step 4:
[1114] The server reformats the processed video data.
[1115] Input: High-quality video data, user-selectable format
[1116] Output: Video file in the specified format
[1117] What happens: The server converts the processed video data into the format selected by the user (e.g. MP4, DVD, etc.) using the appropriate encoding software.
[1118] Step 5:
[1119] The server generates a download link for the video data and provides it to the user.
[1120] Input: Format converted video file
[1121] Output: Download link
[1122] Specific operation: The server hosts the converted video file and generates a download link that users can access. The server notifies users of this link so that they can easily obtain the data.
[1123] Step 6:
[1124] Users can download and use high-quality video files.
[1125] Input: Download link
[1126] Output: Locally saved video file
[1127] Specific operation: Users can click the download link provided to download the high-quality video file to their device and play or share it.
[1128] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1129] The present invention is a system that uses generation AI to convert video data digitized from old videotapes into high-quality image and sound, combines it with an emotion engine that recognizes the user's emotions, processes and reformats the data according to the user's emotions, and finally dubs the data into a format that can be played on DVDs, smartphones, etc. The main modes for implementing the present invention are described in detail below.
[1130] Uploading video data
[1131] Users convert old videotapes (e.g., VHS or miniDV) into digital files using a dedicated digitizing device and save them on their device. Next, they open a specific website or application on their device to upload the saved digital data to a server. The web page provides an interface for selecting and uploading files.
[1132] Example: A user uses a home video capture device to digitize their VHS tapes and save them on their computer. Then, they access a dedicated website, select the saved video files, and upload them.
[1133] High-quality image and sound processing
[1134] The server receives and temporarily stores the uploaded video file. It then processes the received data using a generative AI model. The generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs noise reduction and clearing on the audio track to improve sound quality.
[1135] Example: A generative AI model upscales an original 480p home video to 1080p and produces a noise-reduced video. The audio track is processed to remove background noise and improve audio clarity.
[1136] User Emotion Recognition
[1137] The server analyzes the high-quality video data and invokes an emotion engine to recognize the user's emotions. The emotion engine analyzes the facial expressions and tone of voice of the people in the video data to detect emotions such as smiling or crying. Based on this emotion, filters and effects are added to specific parts of the video.
[1138] Example: An emotion engine analyzes a scene in a home video where a child blows out a birthday cake, and if it detects a happy expression, it adds a special effect to the scene.
[1139] Reformatting and providing data
[1140] Once the processing is complete, the server converts the high-quality video file to the format of the user's choice (e.g. DVD, MP4, etc.). The converted data is then re-saved and a download link is generated. This link is then communicated to the user for easy access.
[1141] The user can then use the provided link to download the converted video data to their device. They can then burn the downloaded file to a DVD, play it on their smartphone, or share the video with family and friends using a communication app or cloud service for data sharing.
[1142] For example, a user downloads an ISO file provided by a server and burns it to a DVD disc using home DVD burning software. Then, the user plays the DVD on a home DVD player and watches it with his or her family. Alternatively, the user uploads the converted MP4 file to cloud storage and shares it with family and friends using a specific communication application.
[1143] Thus, the present invention is a system that uses generative AI and an emotion engine to convert the video and audio from old videotapes into high-quality footage, apply specific processing based on the user's emotions, and make it easy to play and share on modern devices.
[1144] The processing flow will be explained below.
[1145] Step 1:
[1146] Users convert old videotapes (VHS, miniDV, etc.) into digital data using dedicated digitizing equipment. The digitized video data is then stored on the user's device.
[1147] Step 2:
[1148] The terminal provides an interface for uploading the digitized video data to the server through a dedicated website or application. The user selects a file using the terminal interface and clicks the "upload" button.
[1149] Step 3:
[1150] The server receives the video data uploaded by the user and temporarily stores it. When storing it, it analyzes and extracts the metadata of the video data (resolution, frame rate, audio format, etc.).
[1151] Step 4:
[1152] The server then calls the generative AI model, which uses deep learning techniques to enhance the image quality of each frame, remove noise and blur, and perform color correction.
[1153] Step 5:
[1154] The server uses a generative AI model to process the extracted audio track to improve its quality, specifically by removing noise, canceling echoes, and adjusting pitch and clearing to improve sound quality.
[1155] Step 6:
[1156] The server recombines the high-quality video and audio data to generate a new video file, which is then reformatted to the format specified by the user (e.g., DVD format, MP4 format).
[1157] Step 7:
[1158] The server invokes the emotion engine to analyze the facial expressions and voices of people in the video data, and the emotion engine detects specific emotions, such as smiling or crying.
[1159] Step 8:
[1160] The server processes the video by adding filters and effects to specific parts of the video based on the emotions detected by the emotion engine. For example, a special effect is applied to scenes where the emotion of joy is detected.
[1161] Step 9:
[1162] The server reformats the emotion-processed, high-quality video data into a format specified by the user, and generates and saves the final video file.
[1163] Step 10:
[1164] The server stores the generated reformatted video file and makes it accessible to users via a download link, which is communicated to users via email and / or their dashboard.
[1165] Step 11:
[1166] The user can then download the converted video file to their device using the provided download link, then burn it to a DVD or play it on their smartphone.
[1167] Step 12:
[1168] The device provides a communication application for sharing downloaded video files with family and friends. The user can use this application to upload video files to a cloud service or send them directly to family and friends.
[1169] Example 2
[1170] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1171] Due to aging and differences in formats, it is difficult to play the video and audio from old videotapes on modern devices. Furthermore, these videos suffer from noise and low resolution, resulting in a poor viewing experience. Furthermore, there is a lack of technology to recognize user emotions and enrich the visual expression. Therefore, current technology is required to improve the image and sound quality of old videotapes, convert them into formats playable on modern devices, and even add effects that respond to user emotions.
[1172] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for receiving video data digitized from an old video tape, a means for improving the image quality of the digitized video data using a generative AI model, and a means for improving the sound quality of the audio of the digitized video data using a generative AI model. This makes it possible to analyze the video data with high image quality and sound quality and recognize the user's emotions. In addition, by including a means for adding filters and effects to the video data based on the recognized emotions, the visual expression can be enriched.
[1173] "Digitization" is the process of converting analog data into digital data.
[1174] A "generative AI model" is an algorithm that uses artificial intelligence to generate or transform data for a specific task using techniques such as deep learning.
[1175] "High-definition" refers to the process of improving the resolution and image quality of video data.
[1176] "High quality audio" is the process of improving the quality of audio data and removing noise.
[1177] "Emotion recognition" is a technology that analyzes facial expressions and vocal tones of people contained in video data to detect specific emotional states (such as joy, anger, sadness, or happiness).
[1178] A "filter" is a function for applying a specific effect to video data.
[1179] "Effects" is a function that adds special visual effects to video data.
[1180] "Reformatting" is the process of converting video data into a different form or format.
[1181] "Providing" refers to making the generated data available to users.
[1182] "Server" means a central control unit for receiving, storing, processing, and providing video data.
[1183] "User" means the person or end user who utilizes the system to upload video data and download the final product.
[1184] This invention is a system that uses a generative AI model to convert video data digitized from old videotapes into high-quality image and sound, combines it with an emotion engine that recognizes the user's emotions, processes and reformats the data according to the user's emotions, and finally dubs it into a format that can be played on DVDs, smartphones, etc. Specific steps for implementing this invention are described below.
[1185] Video data digitization and uploading
[1186] A user converts old videotapes (e.g., VHS or miniDV) into digital files using dedicated digitizing equipment. This can be done using a home video capture device. The digitized video data is saved on the device. The user then opens a specific website or application on the device and uploads the saved video data to a server. The website provides an interface for selecting and uploading files, and the user selects and uploads the digital data.
[1187] High-quality image and sound processing
[1188] The server receives and temporarily stores video files uploaded by users. The server then launches a generative AI model to process the uploaded video data. This generative AI model uses deep learning techniques to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs similar noise reduction and sound quality clearing on the audio track to improve sound quality. For example, the generative AI model upscales a 480p home video to 1080p and generates a noise-free video. It also performs processing to remove background noise from the audio track and improve sound quality.
[1189] User Emotion Recognition
[1190] The server analyzes the high-quality video data and invokes an emotion engine to recognize the user's emotions. The emotion engine analyzes the facial expressions and tone of voice of the people in the video data to detect emotions such as smiling or crying. Based on these emotions, filters and effects are added to specific parts of the video. For example, if the emotion engine analyzes a scene in a home video where a child blows out a birthday cake and detects a happy expression, it adds a special effect to that scene.
[1191] Reformatting and providing data
[1192] The server converts the processed video file with high image and sound quality into the format specified by the user (e.g., DVD, MP4, etc.). The converted data is then saved again, and a download link is generated and sent to the user. The user then downloads the video data to their device via the link. The user can then burn the downloaded file to a DVD or play it on their smartphone. They can also share the video with family and friends using cloud services or communication applications. For example, a user can download an ISO file provided by the server and burn it to a DVD using their home DVD burning software. They can then play the DVD on a home DVD player and enjoy it with their family. They can also upload the converted MP4 file to cloud storage and share it with family and friends using a specific communication application.
[1193] Prompt Sentence Examples
[1194] 1. "What are the specific steps to digitize and upload my old VHS tapes?"
[1195] 2. "How can I use a generative AI model to convert uploaded videos into high-quality video and audio?"
[1196] 3. "Please explain how you can analyze video data, recognize user emotions, and add specific effects."
[1197] The above is a specific embodiment for implementing the present invention. The present invention is a system for converting old videotape digital data into high-quality data, processing it according to the user's emotions, and making it easy to play and share on modern devices.
[1198] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1199] Step 1: Digitize and upload your video data
[1200] A user converts old videotapes (e.g., VHS or miniDV) into digital files using dedicated digitizing equipment (e.g., a home video capture device). The converted digital video data is stored on the device. Next, the user opens a specific website or application on the device to upload the saved digital data to a server. Through the file upload interface, the user selects the video data and clicks the upload button. The server receives this video data and temporarily stores it. The input is a digitized video file, and the output is video data temporarily stored on the server.
[1201] Step 2: High-quality image and sound processing
[1202] The server retrieves the stored video data and launches a generative AI model. When the video data is input into the generative AI model, it uses deep learning technology to improve the resolution of each frame, remove noise and blur, and perform color correction. It also performs noise reduction and sound quality clearing on the audio track. As a specific example, the generative AI model upscales a 480p home video to 1080p and generates a noise-free video. At the same time, it removes background noise from the audio track and performs processing to clear the sound quality. The input is temporarily stored video data, and the output is video data with high image quality and sound quality.
[1203] Step 3: Recognizing user emotions
[1204] The server receives high-quality video data from the generative AI model and passes it to the emotion engine. When video data is input into the emotion engine, it analyzes the facial expressions and voice tones of the people in the video to detect emotions such as smiling or crying. Based on the emotions recognized by the emotion engine, filters and effects are added to specific scenes. For example, if the emotion engine analyzes a scene of a child blowing out a birthday cake and detects a joyful expression, it adds a special effect to that scene. The input is high-quality video data, and the output is video data with emotion recognition and effects applied.
[1205] Step 4: Reformat and provide data
[1206] Once processing is complete, the server converts the video data, which has undergone high-quality image and sound quality enhancement and emotion recognition processing, into the format specified by the user (e.g., DVD, MP4, etc.). The converted data is then saved back to the server, and a download link is generated. The user then downloads the converted video data to their device via this download link. For example, a user downloads an ISO file provided by the server and burns it to a DVD using their home DVD burning software. The user then plays the DVD on a home DVD player and watches it with their family. Alternatively, the user can upload the converted MP4 file to cloud storage and share it with family and friends using a specific communication application. The input is video data with emotion recognition and effects applied, and the output is video data converted to the user-specified format.
[1207] The above describes the specific operations and data flow for each processing step. This system allows us to convert old videotapes to high quality, add visual effects based on the user's emotions, and provide them in a format that can be played and shared on modern devices.
[1208] (Application example 2)
[1209] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1210] Video and audio stored on old videotapes are not suitable for playback on modern devices with high image and sound quality. Furthermore, digitized data from these videotapes suffers from degradation of the original image and sound quality, making it difficult to achieve visual and auditory satisfaction. Furthermore, the video recorded on home videotapes is filled with people's memories and emotions, and there is a demand for content that reflects those emotions. To solve these issues, not only is it necessary to improve the quality of video data, but it is also necessary to edit it in a way that reflects the user's emotions.
[1211] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1212] In this invention, the server includes means for receiving video data digitized from old videotapes, means for improving the image quality of the digitized video data using a generative AI model, means for improving the audio quality of the digitized video data using a generative AI model, means for reformatting the video data with improved image quality and audio quality, means for providing the reformatted video data to a user, means for recognizing the emotions of people in the video data using an emotion engine and adding effects according to the recognized emotions, and means for outputting the video data in a specified format. This makes it possible to convert data from old videotapes into high-quality image and audio and provide it as video that reflects the user's emotions.
[1213] "Old videotapes" are magnetic tape media that record video and audio in a non-digital analog format, such as VHS or miniDV.
[1214] "Digitalized video data" means video and audio data recorded on analog videotape that has been converted into a digital format that can be processed and stored on an electronic device.
[1215] A "generative AI model" is an artificial intelligence model designed based on deep learning technology, and is an algorithm that analyzes and processes input data to generate specific results.
[1216] "High-definition" refers to the process of processing video data that has been degraded by low resolution or noise using a generative AI model to convert it into high-resolution, clear video.
[1217] "High-quality sound" refers to processing audio data containing background noise and degraded sound using a generative AI model to convert it into clear, high-quality sound.
[1218] "Reformatting" refers to the process of converting digitized video data into a specific file format or designated media format.
[1219] An "emotion engine" is an artificial intelligence technology that analyzes facial expressions and tone of voice of people contained in video and audio data and recognizes their emotions.
[1220] "Effects" refers to the process of adding specific visual or auditory effects to video or audio.
[1221] "Video data metadata" refers to information related to video data, such as supplementary information such as the date and time of filming, the location of filming, and the characters appearing in the video.
[1222] The "specified format" refers to the particular file or media format selected by the user for the final output of the video data.
[1223] This invention is a system that uses generative AI to convert video data digitized from old videotapes into high-quality image and sound, then combines it with an emotion engine that recognizes the user's emotions to add specific effects to the video data, and finally provides it in a format specified by the user.
[1224] Overall system configuration
[1225] The system mainly includes the following components:
[1226] 1. Digitization tools: Video capture devices to generate digital data from old videotapes.
[1227] 2. Data receiving means: A terminal (e.g., a smartphone or PC) that provides an interface for uploading digitized video data to a server.
[1228] 3. Generative AI model: A deep learning model for processing digital video and audio data to produce high-quality images and sounds.
[1229] 4. Emotion Engine: An artificial intelligence engine to analyze the emotions of people in video data and add specific effects.
[1230] 5. Reformatting means: A device that converts video data with high image quality, high sound quality, and added emotional effects into a format specified by the user (e.g., MP4, DVD ISO).
[1231] 6. Data providing means: A server that generates a download link to provide the final video data to the user.
[1232] What the program does
[1233] 1. Upload video data:
[1234] The user starts a dedicated application on a device (such as a smartphone or PC) and uploads the digitized files of old videotapes to the server. At this stage, a file selection interface is provided.
[1235] 2. Processing by generative AI models:
[1236] The server temporarily stores the uploaded video data and uses a generative AI model to improve the video and audio quality. This model is based on deep learning technology using TensorFlow. Specifically, it improves the video resolution, removes noise and blur, and performs color correction. It also adjusts the audio track to remove background noise and ensure clear sound quality.
[1237] 3. Emotion engine processing:
[1238] The video data, which has been enhanced in quality by the generative AI model, is then analyzed using an emotion engine. The emotion engine uses OpenCV and other technologies to analyze the facial expressions and tone of voice of the people in the video, and adds specific effects depending on the emotions it recognizes. For example, in a video of a birthday party, a congratulatory effect can be added to scenes where a child is happy.
[1239] 4. Data Reformatting and Provision:
[1240] The final processed video data is converted into the format specified by the user (e.g. MP4 or DVD ISO). The server saves the converted video data in storage and notifies the user of a download link. Using this link, the user can download the video data and save it on their device.
[1241] Prompt Sentence Examples
[1242] "Convert your 480p home videos to high-quality 1080p footage, remove background noise and improve the clarity of your audio tracks."
[1243] As described above, the present invention is a system that converts digital data from old videotapes into high-quality data and adds effects that correspond to the user's emotions, thereby providing memorable videos in a more attractive form.
[1244] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1245] Step 1:
[1246] Users launch a dedicated application on their device (smartphone or PC), select the digitized files of old videotapes, and upload them to the server.
[1247] Input: Digitized video data file
[1248] Process: The user selects a file using the file selection interface and presses the "Upload" button.
[1249] Output: Video data is uploaded to the server and temporarily stored.
[1250] Step 2:
[1251] The server temporarily stores the uploaded video data and uses generative AI models to improve the video and audio quality.
[1252] Input: Uploaded video data
[1253] Processing: Generative AI models are used to improve video resolution, remove noise and blur, and perform color correction. Specifically, deep learning models using TensorFlow perform frame-by-frame processing. Audio tracks are also adjusted to remove background noise and ensure clarity.
[1254] Output: High-quality video data
[1255] Step 3:
[1256] The emotion engine analyzes high-quality video data and audio to recognize people's emotions.
[1257] Input: High-quality video data
[1258] Processing: An emotion engine (using OpenCV, etc.) is used to analyze facial expressions and vocal tones of people in the video to recognize specific emotions.
[1259] Output: Metadata about the recognized emotion
[1260] Step 4:
[1261] The server adds specific effects to the video data according to the emotions recognized by the emotion engine.
[1262] Input: High-quality video data and emotion recognition metadata
[1263] Processing: Add effects to specific scenes based on the recognized emotion. For example, add a blessing effect to a scene where the emotion is recognized as "joy."
[1264] Output: Video data with effects added
[1265] Step 5:
[1266] The server then converts the final processed video data into the format specified by the user and generates a download link.
[1267] Input: Video data with effects added
[1268] Processing: Convert the data to the desired format (e.g. MP4, DVD ISO) and save it to storage, using a reformatting tool to convert it to the appropriate file format, and generate a link for the user to download it.
[1269] Output: Video data in the specified format and a download link
[1270] Step 6:
[1271] The user uses the provided download link to download the final video data to their terminal.
[1272] Input: Download link
[1273] Processing: The user clicks on the notified link and downloads the video data.
[1274] Output: The final video data saved on your device
[1275] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1276] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1277] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1278] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1279] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1280] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1281] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1282] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1283] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1284] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1285] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1286] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1287] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1288] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1289] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1290] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1291] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1292] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1293] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1294] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1295] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1296] The following is further disclosed regarding the above embodiment.
[1297] (Claim 1)
[1298] means for receiving digitized video data from old videotapes;
[1299] A means for enhancing the quality of the digitized video data using a generative AI model;
[1300] A means for enhancing the audio quality of digitized video data using a generative AI model; and
[1301] means for reformatting the video data with high image quality and sound quality;
[1302] means for providing the reformatted video data to a user;
[1303] A system including:
[1304] (Claim 2)
[1305] 10. The system of claim 1, further comprising: means for extracting metadata from the video data.
[1306] (Claim 3)
[1307] 10. The system of claim 1, further comprising means for converting the obtained high-quality video data into a specified format.
[1308] "Example 1"
[1309] (Claim 1)
[1310] means for receiving digitized video data from an old analog recording medium;
[1311] A means of converting the data into a digital file using specialized digitizing equipment;
[1312] A means for enhancing the quality of the digitized video data using a generative AI model;
[1313] A means for enhancing the audio quality of digitized video data using a generative AI model; and
[1314] means for converting the video data with high image quality and high sound quality into a format designated by a user;
[1315] means for providing the reformatted video data to a user;
[1316] A system including:
[1317] (Claim 2)
[1318] 10. The system of claim 1, further comprising: means for extracting metadata from the video data.
[1319] (Claim 3)
[1320] 10. The system of claim 1, further comprising: means for generating and notifying the designated download link.
[1321] "Application Example 1"
[1322] (Claim 1)
[1323] means for receiving digitized video data from old videotapes;
[1324] A means for enhancing the quality of the digitized video data using a generative AI model;
[1325] A means for enhancing the audio quality of digitized video data using a generative AI model; and
[1326] means for reformatting the video data with high image quality and sound quality;
[1327] means for providing the reformatted video data to a user;
[1328] A means for converting video data into high-quality image and sound on a cloud server and making the video data available for download in a format selected by the user;
[1329] A system including:
[1330] (Claim 2)
[1331] 10. The system of claim 1, further comprising: means for extracting metadata from the video data.
[1332] (Claim 3)
[1333] 10. The system of claim 1, further comprising means for converting the obtained high-quality video data into a specified format.
[1334] "Example 2: Combining Emotion Engines"
[1335] (Claim 1)
[1336] means for receiving digitized video data from old videotapes;
[1337] A means for enhancing the quality of the digitized video data using a generative AI model;
[1338] A means for enhancing the audio quality of digitized video data using a generative AI model; and
[1339] A means for analyzing the video data with high image quality and high sound quality and recognizing the user's emotions;
[1340] means for adding filters or effects to footage of the video data based on the recognized emotions;
[1341] A means for reformatting the video data that has been processed with high image quality, high sound quality, and emotion recognition;
[1342] means for providing the reformatted video data to a user;
[1343] A system including:
[1344] (Claim 2)
[1345] 10. The system of claim 1, further comprising: means for extracting metadata from the video data.
[1346] (Claim 3)
[1347] 10. The system of claim 1, further comprising means for converting the obtained high-quality video data into a specified format.
[1348] "Application example 2 when combining emotion engines"
[1349] (Claim 1)
[1350] means for receiving digitized video data from old videotapes;
[1351] A means for enhancing the quality of the digitized video data using a generative AI model;
[1352] A means for enhancing the audio quality of digitized video data using a generative AI model; and
[1353] means for reformatting the video data with high image quality and sound quality;
[1354] means for providing the reformatted video data to a user;
[1355] a means for recognizing emotions of people in video data using an emotion engine and adding effects according to the recognized emotions;
[1356] means for outputting video data in a specified format;
[1357] A system including:
[1358] (Claim 2)
[1359] 10. The system of claim 1, further comprising: means for extracting metadata from the video data.
[1360] (Claim 3)
[1361] 10. The system of claim 1, further comprising means for converting the obtained high-quality video data into a specified format. [Explanation of symbols]
[1362] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving digitized video data from old videotapes; A means for enhancing the quality of the digitized video data using a generative AI model; A means for enhancing the audio quality of digitized video data using a generative AI model; and means for reformatting the video data with high image quality and sound quality; means for providing the reformatted video data to a user; A system including:
2. The system of claim 1 further comprising: means for extracting metadata from the video data.
3. 2. The system of claim 1, further comprising means for converting the obtained high quality video data into a specified format.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A