System
The system automates the colorization of black-and-white films by preprocessing frames and using a generative AI model to achieve high-quality, efficient, and realistic colorization, addressing the inefficiencies of manual methods.
Patent Information
- Application Number
- JP2024122736
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-02-10
AI Technical Summary
Colorizing traditional black-and-white films and video works requires time-consuming and labor-intensive manual work, making it difficult to realistically recreate the atmosphere and seasonal feel, and is expensive, making it unrealistic to implement on many video works.
A system that includes uploading a video file to a server, extracting and preprocessing frames, colorizing the frames using a generative AI model, reconstructing the colorized frames, and providing a download link for the saved video file, with preprocessing steps including resolution unification and noise removal, and the AI model using learned data for realistic colorization.
Enables realistic colorization of monochrome videos in a shorter time than manual methods, significantly improving visual realism and enhancing the sense of visual appeal of historical and vintage films.
Smart Images

Figure 2026021054000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Colorizing traditional black-and-white films and video works requires time-consuming and labor-intensive manual work. This makes it difficult to realistically recreate the atmosphere and seasonal feel of the time, and is insufficient to convey the original intent to viewers. Furthermore, manual colorization is expensive, making it unrealistic to implement it on many video works. Therefore, a method is needed to solve these issues and achieve high-quality colorization in a short amount of time. [Means for solving the problem]
[0005] The present invention is a system that includes a means for uploading a video file to a server, a means for extracting and preprocessing frames from the video file, a means for colorizing the preprocessed frames using a generative AI model, a means for reconstructing the colorized frames and saving them as a video file, and a means for providing a user with a download link for the saved video file.
[0006] Specifically, the preprocessing means includes means for unifying the resolution of each frame and performing noise removal and contrast adjustment. The generative AI model also includes means for coloring each frame based on the learned data. This makes it possible to provide realistic color images in a shorter time than manual colorization, significantly improving the sense of visual realism.
[0007] A "video file" is a digital file containing video data, and refers to various video content such as movies, television programs, and documentaries.
[0008] A "server" refers to a computer system that provides services to other computers (clients) on a computer network.
[0009] A "frame" refers to an individual still image that makes up a video, and a video is formed by a series of these in chronological order.
[0010] "Preprocessing" refers to initial processing carried out to improve the quality of video data, and specifically refers to operations such as unifying resolution, removing noise, and adjusting contrast.
[0011] A "generative AI model" is a model trained using machine learning algorithms that has the ability to generate new data based on input data or transform input data.
[0012] "Colorization" refers to the process of adding color information to monochrome video or images to convert them into color video or images.
[0013] "Reconstruction" refers to the reassembly of divided or processed data into a single data format, and includes the operation of reassembling video frames into a video.
[0014] A "download link" is a URL (Uniform Resource Locator) for obtaining a file on the Internet, and by accessing it, a user can save specific data on their device. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] The present invention is a system for automatically colorizing monochrome movies and video works, and has the following procedures and functions: A user, a server, and a terminal work together.
[0037] System Overview
[0038] 1. Uploading video files
[0039] The user uploads monochrome video files to the server through a specified interface from their own terminal. The user accesses the server's upload page using a web browser, selects a local video file using a file selection dialog, and performs the upload operation.
[0040] 2. Frame extraction and preprocessing
[0041] When the server receives the uploaded video file, it first divides it into frames. Each frame undergoes pre-processing to unify the resolution, remove noise, and adjust the contrast, making it suitable for subsequent AI processing.
[0042] 3. Colorization
[0043] The pre-processed frames are input into a generative AI model by the server. The AI model applies appropriate colors to each frame based on pre-trained data. The model achieves realistic colorization using a variety of color data learned from historical documents and literature.
[0044] 4. Reconstructing the frame
[0045] Once the server has captured the colorized frames, it reconstructs them into their original video format, resulting in a colorized video file. The reconstruction process also integrates audio data if necessary.
[0046] 5. Providing a download link
[0047] The server stores the reconstructed color video file and generates a corresponding download link, which is provided to the user, who uses it to download the colorized video file to their device.
[0048] Specific examples
[0049] For example, consider a scenario in which a user wants to upload a monochrome video entitled "Life in Ancient Rome." The user accesses the server's upload page through a web browser on their own device and uploads the video file.
[0050] The server receives the "Life in Ancient Rome" video, splits it into frames, standardizes the resolution of each frame to 720p, removes noise, and adjusts the contrast, preparing the pre-processed frames.
[0051] The frames are then fed into a generative AI model on the server, which then colors each frame with the colors most appropriate for it, based on historical documents about the colors of Roman architecture and clothing, for example.
[0052] The colorized frames are then reassembled and saved as a single video file. Finally, users can use the download link provided by the server to download the colorized version of "Life in Ancient Rome" to their device and watch it.
[0053] This system enables the realistic colorization of black-and-white footage in a short amount of time, and is expected to significantly enhance the visual value of historical film works.
[0054] The processing flow will be explained below.
[0055] Step 1:
[0056] Users upload monochrome video files to the server from their own devices through a designated interface. This operation is performed using a web browser, accessing the upload page, selecting a local video file from the file selection dialog, and pressing the upload button.
[0057] Step 2:
[0058] The server receives the uploaded video file, splits the video data into frames, and uses a video processing library to extract each frame and temporarily store it in memory.
[0059] Step 3:
[0060] The server performs pre-processing on each extracted frame. First, it resizes it to the specified resolution (e.g., 720p) to ensure uniformity. Then it applies a noise reduction filter to improve the image quality. It also performs contrast adjustments to ensure visual consistency across the frames.
[0061] Step 4:
[0062] The server feeds the pre-processed frames into a trained generative AI model, which runs an algorithm to apply the appropriate color to each frame. The model recognizes certain features and, based on those, selects the appropriate color to apply to the frame.
[0063] Step 5:
[0064] The server then reconstructs the colorized frames into a single video file by using a video processing library to sequence the frames, combine them with the original audio data, and save the resulting video file.
[0065] Step 6:
[0066] The server saves the completed colorized video file in the server, generates a link so that the user can download it, and notifies the user of the generated link.
[0067] Step 7:
[0068] Users can download the colorized video file to their device using the download link provided by the server. When they click the link, the file is automatically saved in the download folder.
[0069] As a concrete example, if a user uploads a black-and-white video titled "Life in Ancient Rome," the video will be colorized through the steps described above. The user can then download the colorized version of "Life in Ancient Rome" and view it on their device. This process makes it possible to provide high-quality colorized video in a short amount of time.
[0070] Example 1
[0071] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0072] Conventional techniques for colorizing black-and-white footage require significant time and effort, requiring manual work by people with specialized skills and knowledge. Furthermore, the quality of the colorization is often inconsistent, resulting in unnatural coloring. This makes it difficult to unlock the value of historical and vintage films. Therefore, there is a need for an efficient system for automatically and realistically colorizing black-and-white footage.
[0073] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0074] In this invention, the server includes a means for uploading video files, a means for extracting and preprocessing frames from the video files, a means for colorizing the preprocessed frames using a generative AI model, a means for reconstructing the colorized frames and saving them as video files, and a means for providing users with a download link for the saved video files, thereby enabling realistic colorization of monochrome videos in a short amount of time.
[0075] "Video file" refers to video data stored in digital format.
[0076] "Server" refers to a computer system that provides services over a network.
[0077] "Uploading" refers to the act of transferring data from a local device to a server.
[0078] "Frame" refers to the individual still images that make up a video file.
[0079] "Preprocessing" refers to the preliminary preparation work carried out to improve data quality and for analysis.
[0080] A "generative AI model" refers to a program created using machine learning algorithms that learns patterns from data and generates new data.
[0081] "Colorization" refers to the process of adding color to black and white video or images.
[0082] "Reconstructing" refers to reorganizing the processed data into a new format.
[0083] "Storing" refers to recording and maintaining digital data in a storage device.
[0084] "Download link" refers to a URL for obtaining a digital file over a network.
[0085] "Unifying the resolution" refers to the process of unifying the number of pixels in an image or video according to a certain standard.
[0086] "Noise removal" refers to the process of removing unnecessary random fluctuations from image or video data.
[0087] "Contrast adjustment" refers to the process of adjusting the difference between light and dark to improve the visibility of images and videos.
[0088] "Pre-trained data" refers to an existing dataset that a machine learning model has used during its training process.
[0089] "Coloring" refers to adding appropriate color information to a black and white frame.
[0090] "Historical materials" refers to records and documents about past events and cultures.
[0091] "Realistic colors" refers to realistic and natural color representation.
[0092] This invention is a system for automatically colorizing monochrome movies and video works. This system operates by linking the user's device, a server, and a generative AI model. Specifically, it has the following steps and functions:
[0093] First, the user accesses the server's upload page through a web browser on their device. There, they select a monochrome video file stored on their local disk and upload it to the server. For example, they use the web browser's "file selection" dialog. The file selected by the user is sent to the server via an HTTP POST request.
[0094] When the server receives the uploaded video file, it uses the video processing library FFmpeg to split the video file into frames. Each frame undergoes preprocessing and then the following steps:
[0095] 1. Resolution unification: Resize the resolution of each frame to 720p.
[0096] 2. Noise Reduction: Use a Gaussian filter to remove noise in the video.
[0097] 3. Contrast Adjustment: Perform histogram equalization and adjust the contrast of the frame.
[0098] Once preprocessing is complete, the colorization process is performed using a generative AI model. The server inputs the preprocessed frames into the AI model, which then predicts and applies the optimal color for each frame. The generative AI model used is built using TensorFlow and PyTorch and is trained on a variety of data based on historical documents and literature. For example, based on the training data that "Roman buildings use yellowish stone," a yellowish color is applied to a building.
[0099] The colorized frames are reconstructed again using FFmpeg, and all frames are joined together to form a continuous video file. At this time, if there is audio data from the original video, this is also integrated at the same time. This completes the colorized video file. Specifically, the following FFmpeg command is used to join the frames:
[0100] ffmpeg -r 24 -i frame%04d_colored.png -i original_audio.mp3 -c:v libx264 -c:a aac output_colored_video.mp4
[0101] Finally, the server stores the reconstructed color video file and generates a downloadable link that can be provided to the user, who can use it to download the colorized video file to their device and view it.
[0102] As a concrete example, consider the case where a user uploads a monochrome video of "life in ancient Rome" to a server. The user accesses the server's upload page using a web browser on their device and uploads the "life in ancient Rome" video file. The server then divides the video into frames and preprocesses each frame. The preprocessed frames are then fed into a generative AI model for colorization. For example, each frame is colored with the optimal color based on historical literature on the colors of Roman architecture and clothing. The colorized frames are then reconstructed to generate the final color video file. The user can then download the colorized video file to their device using the provided download link and watch it. This system is expected to rapidly achieve realistic colorization of monochrome video, significantly improving the visual value of historical video works.
[0103] The above is a specific embodiment of the present invention, and the operation and configuration of each part of the system have been specifically described. This system enables realistic and efficient colorization of monochrome footage, significantly reducing the time required compared to traditional manual colorization. Furthermore, the automated process ensures consistently high colorization quality, significantly enhancing the visual appeal of historical footage and vintage films.
[0104] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0105] Step 1:
[0106] The user launches a web browser on their device and accesses the server's upload page. Specifically, they enter the server's upload page URL (e.g., https: / / example.com / upload) in the browser's URL bar. When the page appears, they click the "Choose File" button, select a monochrome video file (e.g., roman_life.mp4) from their local disk, and press the "Upload" button. The input data is a monochrome video file, and the output is a video file uploaded to the server.
[0107] Step 2:
[0108] The server receives the uploaded video file. Then, it splits the video file into frames using FFmpeg. Specifically, it executes the following command:
[0109] ffmpeg -i roman_life.mp4 frame%04d.png
[0110] As a result, the input data is a video file, and the output data is the divided frame images (e.g., frame0001.png, frame0002.png, ...).
[0111] Step 3:
[0112] The server performs pre-processing for each divided frame, which includes the following specific steps:
[0113] Resolution uniformity: To resize each frame to 720p, for example, run the following command:
[0114] ffmpeg -i frame%04d.png -vf scale=1280:720 frame%04d_resized.png
[0115] Denoising: To remove noise using a Gaussian filter, run the following command:
[0116] ffmpeg -i frame%04d_resized.png -vf "gblur=sigma=2" frame%04d_denoised.png
[0117] Contrast adjustment: To perform histogram equalization, run the following command:
[0118] ffmpeg -i frame%04d_denoised.png -vf "eq=contrast=1.5" frame%04d_processed.png
[0119] The input data is a divided frame image, and the output data is a frame image that has been preprocessed.
[0120] Step 4:
[0121] The server inputs the preprocessed frames into a generative AI model for colorization. Specifically, a generative AI model built with TensorFlow or PyTorch is used. The prompt input to the model is in the form of "Please apply appropriate colors based on the frame image I will provide." The input data is the preprocessed frame image, and the output data is the colorized frame image.
[0122] Step 5:
[0123] The server then reconstructs the colorized frames and combines them into a single video file by running the following FFmpeg command:
[0124] ffmpeg -r 24 -i frame%04d_colored.png -i original_audio.mp3 -c:v libx264 -c:a aac output_colored_video.mp4
[0125] The input data are colorized frame images and the original audio data, and the output data is a reconstructed color video file.
[0126] Step 6:
[0127] The server stores the reconstructed color video file and generates a download link, which is provided to the user, who uses it to download the colorized video file to their device. The input data is the reconstructed color video file, and the output data is the download link.
[0128] (Application example 1)
[0129] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0130] Today, there is a growing need to colorize monochrome films and photographs to breathe new life into them. However, traditional colorization methods are time-consuming, costly, and often require specialized knowledge. This creates a need for a system that allows anyone to easily convert monochrome images to color. There is also a need for this functionality to be easily accessible anywhere via mobile devices such as smartphones and tablets.
[0131] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0132] In this invention, the server includes means for uploading video files to the server, means for extracting and preprocessing frames from the video files, means for colorizing the preprocessed frames using a generative AI model, means for reconstructing the colorized frames and saving them as video files, means for providing users with download links for the saved video files, and means for uploading the video files via a smartphone application and making the results available for download. This allows users to easily colorize monochrome videos using their smartphones and quickly obtain the results.
[0133] A "video file" is a digital file containing visual and audio information.
[0134] A "server" is a dedicated computer system that provides services to other computers over a computer network.
[0135] A "frame" is a unit of still images that make up a video file.
[0136] "Preprocessing" refers to processing performed before data analysis or processing, and is an operation performed to improve the quality of the data.
[0137] A "generative AI model" is a trained computational model that uses artificial intelligence technology to analyze data and output results.
[0138] "Colorization" is the process of adding color to monochrome images or videos.
[0139] "Reconstruction" is the process of reassembling decomposed or preprocessed data back into its original form.
[0140] A "download link" is a URL that allows a user to obtain a file via the Internet.
[0141] "User" means a person or entity that uses the system or service.
[0142] A "smartphone application" is a software program that runs on a mobile device.
[0143] This invention is a system for automatically colorizing monochrome video works, and operates in cooperation with a user, a server, and a terminal. Specific embodiments of the invention are described below.
[0144] First, the user launches the application on their smartphone and selects the monochrome video file they want to colorize. This video file is then uploaded to the server via the smartphone application. The uploaded video file is then saved on the server, and the next process begins.
[0145] When the server receives the video file, it divides it into frames and preprocesses each frame. This preprocessing includes unifying the resolution, removing noise, and adjusting the contrast. This is done using software libraries such as OpenCV. Each preprocessed frame is then colorized using a generative AI model. This generative AI model has been trained in advance on a large amount of data, enabling highly accurate colorization.
[0146] The colorized frames are then reconstructed on the server and converted back into the original video file format. The reconstructed color video file is stored on the server, and a download link is generated that users can access. By providing this link to users, they can download the colorized video file via their smartphone.
[0147] Through this process, users can easily colorize monochrome images using their smartphones and receive the results quickly.
[0148] As a concrete example, consider a scenario in which a user uploads a monochrome video titled "Video of an Ancient City." The user first selects and uploads the video file through a smartphone application. The server receives the video file, splits it into frames using OpenCV, standardizes the resolution, removes noise, and adjusts the contrast. Then, a generative AI model colorizes the video, assigning appropriate colors to each frame. The server then reconstructs the colorized frames and generates a download link to provide to the user. Using this link, the user can download and watch the colorized video.
[0149] Another example of a prompt is, "Please input a monochrome frame and generate appropriate colors based on that frame. For example, the building should be the color of stone, the sky should be blue, and the ground should be brown," and the AI model will generate appropriate colors.
[0150] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0151] Step 1:
[0152] A user launches a smartphone application, selects a monochrome video file, and uploads it. This operation causes the smartphone application to send the selected video file to the server. The input is the user's video file, and the output is the file uploaded to the server. At this time, an HTTP POST request is used to send the file to the server.
[0153] Step 2:
[0154] The server receives the uploaded video file and splits it into frames. The input is the video file uploaded to the server, and the output is the split frames. The server uses the OpenCV library to perform the specific operation of splitting the video file into frames.
[0155] Step 3:
[0156] The server preprocesses each frame. The input is a set of split frames, and the output is a set of preprocessed frames. Preprocessing includes unifying the resolution, removing noise, and adjusting the contrast. Specifically, OpenCV is used to unify the resolution of each frame to 720p, apply a noise removal filter, and adjust the contrast using histogram equalization.
[0157] Step 4:
[0158] The server inputs the preprocessed frames into a generative AI model for colorization. The input is a set of preprocessed frames, and the output is a set of colorized frames. The generative AI model applies the appropriate color to each frame based on pre-trained data. Specifically, it uses TensorFlow or PyTorch models to colorize each frame.
[0159] Step 5:
[0160] The server reconstructs the colorized frames and returns them to the original video file format. The input is a set of colorized frames, and the output is a reconstructed color video file. The server uses libraries such as OpenCV to perform the specific operations of successively combining each frame and reconstructing it into the original video format.
[0161] Step 6:
[0162] The server saves the reconstructed color video file and generates a download link for users to access. The input is the reconstructed color video file, and the output is the download link. The server then performs the specific operation of generating a URL for the saved video file and providing it to the user.
[0163] Step 7:
[0164] A user accesses a download link through a smartphone application and downloads a colorized video file. The input is the download link, and the output is the colorized video file on the smartphone. The user clicks on the generated link and performs the specific action of downloading the file using an HTTP GET request.
[0165] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0166] The present invention relates to a system that automatically colorizes monochrome movies and video works by combining an emotion engine with a system that adjusts colors and filters based on the user's emotions. Specific embodiments of the present invention will be described below.
[0167] System Overview
[0168] 1. Uploading video files
[0169] The user uploads monochrome video files to the server through a specified interface from their own terminal. The user accesses the server's upload page using a web browser, selects a local video file using a file selection dialog, and performs the upload operation.
[0170] 2. Frame extraction and preprocessing
[0171] When the server receives the uploaded video file, it divides the video data into frames. Each frame undergoes preprocessing to unify the resolution, remove noise, and adjust the contrast, making it suitable for subsequent AI processing.
[0172] 3. Analysis by Emotion Engine
[0173] An emotion engine installed on the server recognizes the user's emotions. The emotion engine collects the user's facial expression data, voice data, or biometric information, and analyzes this data to identify the user's emotions. For example, if the user shows a happy expression while watching, that emotion data is analyzed.
[0174] 4. Adjusting the colorization process
[0175] Based on the analysis results of the emotion engine, the colorization parameters for the preprocessed frames are adjusted. For example, if the user is emotional, warm colors are emphasized. The preprocessed frames are input into the generative AI model, which then colorizes them based on the adjusted parameters.
[0176] 5. Reconstructing the frame
[0177] The server then reassembles the colorized frames into a single video file, resulting in a complete colorized video file. The reconstruction process also integrates audio data, if necessary.
[0178] 6. Providing download links
[0179] The server stores the reconstructed color video file in the server and generates a link for the user to download it. The created link is notified to the user, who uses it to download the colorized video file to their terminal.
[0180] Specific examples
[0181] For example, consider a scenario in which a user uploads a black and white video entitled "Life in Ancient Rome." The user accesses the server's upload page through a web browser and uploads this video file.
[0182] The server receives the "Life in Ancient Rome" video and splits it into frames, then standardizes the resolution of each frame to 720p, denoises it, and adjusts the contrast, preparing the pre-processed frames.
[0183] If the emotion engine analyzes the user's facial expressions and voice and recognizes that the user has positive emotions toward the video, it adjusts the colorization parameters based on this emotion data to emphasize warm colors.
[0184] The pre-processed frames are then fed into a generative AI model, which then colorizes them based on the adjusted parameters. For example, each frame is colored with the optimal color based on historical documents about the colors of Roman architecture and clothing.
[0185] The colorized frames are then reassembled and saved as a single video file. Finally, users can download the colorized version of "Life in Ancient Rome" to their devices using a download link provided by the server.
[0186] The present invention is expected to provide realistic colorized images based on the user's emotions in a short time, significantly improving visual value and user experience.
[0187] The processing flow will be explained below.
[0188] Step 1:
[0189] Users upload monochrome video files to the server from their own devices through a designated interface. This operation is performed using a web browser, accessing the upload page, selecting a local video file from the file selection dialog, and pressing the upload button.
[0190] Step 2:
[0191] The server receives the uploaded video file and temporarily stores it in a specified directory on the server, before moving on to the next processing step.
[0192] Step 3:
[0193] The server splits the video file into frames, analyzes the video frame by frame using a video processing library, extracts each frame as an image file, and stores it in temporary storage.
[0194] Step 4:
[0195] The server performs pre-processing on each extracted frame. First, it resizes it to the specified resolution (e.g., 720p) to ensure uniformity. Then it applies a noise reduction filter to improve the image quality. It also performs contrast adjustments to ensure visual consistency across the frames.
[0196] Step 5:
[0197] The emotion engine installed on the server recognizes the user's emotions. It collects facial expressions and voice data from the video the user is watching, and analyzes this data to identify the user's emotions. The emotion engine uses facial recognition and voice analysis technology to determine the user's emotional state (happiness, surprise, sadness, etc.).
[0198] Step 6:
[0199] The server automatically adjusts the colorization parameters based on the analysis results of the emotion engine. For example, if the user is recognized as happy, the colors are set to emphasize warm colors. The parameters provided by the emotion engine become input data for the generative AI model.
[0200] Step 7:
[0201] The server inputs the preprocessed frames into the generative AI model based on the adjusted parameters. The AI model selects the optimal color for each frame and colorizes it. The model then references the trained dataset and runs an algorithm to generate realistic colors.
[0202] Step 8:
[0203] The server reconstructs the colorized frames into a single video file, sequences the colorized frames using a video processing library, combines them with the original audio data, and saves the resulting video file.
[0204] Step 9:
[0205] The server stores the generated color image file in the server and generates a link for the user to download it, and notifies the user of the generated link.
[0206] Step 10:
[0207] Users can download the colorized video file to their device using the download link provided by the server. When they click the link, the file is automatically saved in the download folder.
[0208] Specific examples
[0209] For example, if a user uploads a monochrome video titled "Life in Ancient Rome," the video will be colorized through the steps described above. If the emotion engine analyzes the user's facial expressions and voice data and determines that the user has a positive emotion toward the video, the colorization will be adjusted to emphasize warm colors. This process allows the user to download and watch "Life in Ancient Rome," with enhanced visual value and emotional empathy. The present invention is expected to provide realistic colorized video tailored to the emotions of individual users in a short time, significantly improving visual value and user experience.
[0210] Example 2
[0211] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0212] Conventional colorization technologies simply colorize monochrome images and are unable to reflect the user's emotions or visual preferences. As a result, there is a gap between the colorized image and the user's emotions, resulting in a degradation of the viewing experience. Furthermore, due to inconsistent colorization accuracy and quality, parts of the image can appear unnatural.
[0213] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0214] In this invention, the server includes means for uploading a video file to an information processing device, means for extracting and preprocessing frames from the video file, means for collecting and analyzing user emotional information, means for adjusting colorization parameters for the preprocessed frames based on the analysis results, means for colorizing the preprocessed frames using a generative AI model, means for reconstructing the colorized frames and saving them as video files, and means for providing the user with a download link for the saved video file.
[0215] This allows for the provision of more realistic and moving colorized images based on the user's emotions, improving the quality of the viewing experience.
[0216] An "information processing device" is a device for collecting, analyzing, storing, and transmitting data, and includes a server or computer system.
[0217] "Video file" is a digital file containing video data, including monochrome video uploaded by users and final colorized video files.
[0218] A "frame" is a unit of still images that make up a video, and is the basic unit into which a video file is divided when it is processed.
[0219] "Preprocessing" refers to the processing performed on video frames, including resolution unification, noise removal, contrast adjustment, etc.
[0220] "Emotion information" is data that indicates the user's emotional state, and includes facial expression data, voice data, biometric information, and the like.
[0221] An "emotion engine" is a system that analyzes a user's emotional information and identifies their emotional state.
[0222] "Colorization parameters" are setting values for adjusting the color of the frame, and are determined based on the emotion analysis results.
[0223] A "generative AI model" is an artificial intelligence model that colorizes video frames based on pre-trained data.
[0224] "Reconstruction" is the process of connecting the colorized frames in order to recreate the original video file.
[0225] "Download link" refers to a URL or hyperlink that allows a user to download a stored video file via the Internet.
[0226] "Noise reduction" is the process of removing unwanted noise from a video frame.
[0227] "Contrast adjustment" is the process of improving the visibility of an image by adjusting the brightness of the frame.
[0228] "Upload" is the act of a user sending data from their own terminal to a server.
[0229] The present invention relates to a system that automatically colorizes monochrome movies and video works by combining an emotion engine with a system that adjusts colors and filters based on the user's emotions. Specific embodiments of the present invention will be described below.
[0230] (1. Uploading video files)
[0231] The user accesses the upload page of the information processing device using a web browser. The user clicks the "Select File" button on the page and selects a monochrome video file from the local file system. Next, the user presses the upload button to send the video file to the server. The server stores the received video file in a database.
[0232] (2. Frame Extraction and Preprocessing)
[0233] The server passes the received video file to the video analysis module, which divides the video file into frames and sends each frame to the image processing module, which standardizes the resolution to 720p and performs preprocessing such as noise reduction and contrast adjustment. This generates frames suitable for subsequent processing.
[0234] (3. Analysis by Emotion Engine)
[0235] When viewing the video, the user allows the device's camera and microphone to be used. The server collects the user's facial expression data, voice data, and biometric information (e.g., heart rate) in real time. This data is passed to the emotion engine, which analyzes the user's emotions. For example, if the user smiles, the emotion engine analyzes it as "joy."
[0236] (4. Adjusting colorization processing)
[0237] The server's colorization module receives the emotion analysis results from the emotion engine and adjusts the colorization parameters for the preprocessed frames. For example, if the user is emotional, the colorization parameters are set to emphasize warm colors. These adjusted parameters are input as prompts to the generative AI model to colorize each frame. An example prompt is: "The user is emotional. Please use warmer colors."
[0238] (5. Reconstructing the frame)
[0239] The server then reconstructs the colorized frames into a single video file. The reconstruction module then concatenates the frames together to form a continuous video file. If necessary, audio data is also integrated at this stage.
[0240] (6. Providing a download link)
[0241] The server saves the reconstructed color video file in its own storage. It then automatically generates a link that allows the user to download the video file. The user is notified of this link, and can click it to download the colorized video file to their device.
[0242] Specific examples
[0243] For example, consider a scenario in which a user wants to upload a black-and-white video entitled "Life in Ancient Rome." The user uses a web browser to access the upload page of the information processing device and uploads this video file.
[0244] The server receives the "Life in Ancient Rome" video and splits it into frames. Then it standardizes the resolution of each frame to 720p, removes noise, and adjusts the contrast. The pre-processed frames are prepared.
[0245] If the emotion engine analyzes the user's facial expressions and voice and determines that the user has positive emotions toward the video, it adjusts the colorization parameters based on this emotion data to emphasize warm colors. Example prompt: "The user has positive emotions. Use warm colors."
[0246] The pre-processed frames are then fed into a generative AI model, which then colorizes them based on the adjusted parameters. For example, each frame is colored with the optimal color based on historical documents about the colors of Roman architecture and clothing.
[0247] The colorized frames are then reassembled and saved as a single video file. Finally, users can download the colorized version of "Life in Ancient Rome" to their devices using a download link provided by the server.
[0248] The present invention is expected to provide realistic colorized images based on the user's emotions, significantly improving visual satisfaction and user experience.
[0249] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0250] Step 1:
[0251] The user launches a web browser and accesses the upload page of the information processing device. The user clicks the "Select File" button and selects a monochrome video file from the local file system. The user then presses the upload button to send the video file to the server. The input is the video file on the user's local system, and the output is the video file saved on the server.
[0252] Step 2:
[0253] The server passes the received video file to the video analysis module. The video analysis module divides the video file into frames. Each frame is then sent to the image processing module. The image processing module standardizes the resolution to 720p and performs preprocessing such as noise reduction and contrast adjustment. The input is the uploaded video file, and the output is a set of preprocessed frames.
[0254] Step 3:
[0255] When viewing a video, the user allows the device's camera and microphone to be used. The server collects the user's facial expression data, voice data, and biometric information (e.g., heart rate) in real time. This data is passed to the emotion engine. The emotion engine analyzes the user's emotions and generates the analysis results. The input is the user's real-time biometric data, and the output is the user's emotion analysis results.
[0256] Step 4:
[0257] The server's colorization module receives the emotion analysis results from the emotion engine and adjusts the colorization parameters for the preprocessed frames. The emotion analysis results are provided as input, and a prompt sentence is generated based on the results. This prompt sentence is input into a generative AI model to colorize each frame. For example, if the user is emotional, the colorization parameters are set to emphasize warm colors. The input is the emotion analysis results, and the output is the prompt sentence with the adjusted colorization parameters.
[0258] Step 5:
[0259] The server uses a generative AI model to colorize the preprocessed frames using adjusted parameters. Specifically, the generative AI model adds color to each frame based on the prompt. For example, each frame is assigned an appropriate color based on historical documents about the colors of Roman architecture and clothing. The input is the preprocessed frames and the prompt, and the output is a set of colorized frames.
[0260] Step 6:
[0261] The server then reconstructs the colorized frames into a single video file. The reconstruction module concatenates the frames together to form a continuous video file. If necessary, audio data is also integrated at this stage. The input is a set of colorized frames, and the output is a reconstructed color video file.
[0262] Step 7:
[0263] The server saves the reconstructed color video file in its own storage. It then automatically generates a link that allows the user to download the video file. This link is notified to the user, who clicks the link to download the colorized video file to their device. The input is the reconstructed color video file, and the output is the download link.
[0264] (Application example 2)
[0265] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0266] Conventional systems for colorizing black-and-white movies and video works simply add color, and do not provide a personalized viewing experience based on the user's emotions. As a result, there is a problem that the viewing experience is limited because the color adjustment does not correspond to the viewer's emotions. To solve this problem, a system that can colorize in accordance with the user's emotions is needed.
[0267] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading a video file to the information processing device, means for extracting and preprocessing frames from the video file, means for colorizing the preprocessed frames using a generative AI model, means for analyzing a user's emotions, means for adjusting colorization parameters based on the analysis results, means for reconstructing the colorized frames and saving them as a video file, and means for providing a download link for the saved video file. This enables real-time colorization based on the user's emotions.
[0268] "Video file" means a data file that stores the content of visual media in digital format.
[0269] An "information processing device" is a device that processes and manages data, and includes computer systems such as servers.
[0270] A "frame" is a unit of each still image that makes up a video.
[0271] "Preprocessing" refers to the process of preparing data for subsequent processing.
[0272] A "generative AI model" is a model trained using artificial intelligence to perform a specific task, in this case colorization.
[0273] "Means for analyzing emotions" refers to a processing method for identifying emotions based on the user's facial expressions, voice, biometric information, etc.
[0274] A "parameter" is a setting that adjusts the behavior of a system or model.
[0275] "Reconstruction" means reassembling divided data into a single piece of data.
[0276] "Download link" refers to the access means by which a user can save the required data on their device via the Internet.
[0277] The present invention relates to a system for automatically colorizing monochrome movies and video works, and adjusting colors and filters based on the user's emotions. This system is configured and operates as follows.
[0278] First, a user uploads a monochrome video file to the information processing device through a specified interface on their own terminal. The user accesses the upload page using a web browser, selects a local video file using a file selection dialog, and performs the upload operation.
[0279] Next, when the server receives the uploaded video file, it divides the video data into frames. Each frame undergoes preprocessing to unify the resolution, remove noise, and adjust the contrast, making it suitable for subsequent AI processing.
[0280] The preprocessed frames are then analyzed by an emotion analysis engine installed on the server to determine the user's emotional state. The emotion analysis uses the user's facial expression data, voice data, or biometric information. For example, the emotion analysis engine uses cloud services such as EmotionAPI or Amazon Rekognition.
[0281] Based on the analysis, the server generates colorization parameters for the preprocessed frames. A generative AI model (such as OpenAI's GPT-4 or DALL-E) is used to colorize each frame based on these parameters. For example, if the user shows a happy expression, warm colors will be emphasized.
[0282] The colorized frames are then reconstructed into a single video file. During the reconstruction process, audio data is also integrated if necessary. Finally, the colorized video file is stored on a server, and a link is generated so that the user can download it. The created download link is notified to the user, who can use it to download the colorized video file to their device and view it.
[0283] Below are some example prompts that can be used as input to a generative AI model:
[0284] prompt:
[0285] Colorizing black and white "Classic Movies" while analyzing the user's emotional state:
[0286] If the user's emotion is "joy," use warm colors (e.g., warm orange tones)
[0287] If the user is emotional, use soft pastel colors.
[0288] Match the colors of each scene to their historical equivalents
[0289] Detailed frame-by-frame colorization
[0290] This makes it possible to provide an optimal video experience that suits the user's emotions.
[0291] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0292] Step 1:
[0293] A user uploads a monochrome video file to the server through a web browser. The user accesses a web page, selects a video file from a local file selection dialog, and uploads it. The input of this operation is a local monochrome video file, and the output is a video file saved on the server.
[0294] Step 2:
[0295] The server splits the uploaded video file into frames. The server analyzes the video file and extracts each frame. The input is the video file, and the output is a collection of individual frames.
[0296] Step 3:
[0297] The server pre-processes each frame, unifying the resolution of each frame, denoising, and adjusting the contrast. The input is a set of frames, and the output is a set of pre-processed frames.
[0298] Step 4:
[0299] The server analyzes the user's emotions from the preprocessed frames. Using an emotion analysis engine, the server collects and analyzes the user's facial expression data, voice data, and biometric information. The input is the user's facial expression data, voice data, and biometric information, and the output is the user's emotional state.
[0300] Step 5:
[0301] The server generates colorization parameters based on the analysis results. Based on the generated emotion data, the generative AI model generates a prompt sentence, and sets the colorization parameters for each frame based on that prompt sentence. The input is the user's emotional state, and the output is the colorization parameters.
[0302] Step 6:
[0303] The server inputs the preprocessed frames into a generative AI model to perform colorization. The generative AI model (e.g., GPT-4 or DALL-E) colorizes each frame based on the adjusted parameters. The input is a set of preprocessed frames and colorization parameters, and the output is a set of colorized frames.
[0304] Step 7:
[0305] The server reconstructs the colorized frames and saves them as the final color video file. It performs inter-frame reconstruction and integrates audio data as needed. The input is a set of colorized frames and the audio data from the original video, and the output is a reconstructed color video file.
[0306] Step 8:
[0307] The server generates a download link for the reconstructed color video file and provides it to the user. The download link is sent to the user via email or notification. The input is the reconstructed color video file, and the output is the download link provided to the user.
[0308] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0309] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0310] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0311] [Second embodiment]
[0312] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0313] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0314] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0315] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0316] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0317] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0318] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0319] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0320] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0321] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0322] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0323] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0324] The present invention is a system for automatically colorizing monochrome movies and video works, and has the following procedures and functions: A user, a server, and a terminal work together.
[0325] System Overview
[0326] 1. Uploading video files
[0327] The user uploads monochrome video files to the server through a specified interface from their own terminal. The user accesses the server's upload page using a web browser, selects a local video file using a file selection dialog, and performs the upload operation.
[0328] 2. Frame extraction and preprocessing
[0329] When the server receives the uploaded video file, it first divides it into frames. Each frame undergoes pre-processing to unify the resolution, remove noise, and adjust the contrast, making it suitable for subsequent AI processing.
[0330] 3. Colorization
[0331] The pre-processed frames are input into a generative AI model by the server. The AI model applies appropriate colors to each frame based on pre-trained data. The model achieves realistic colorization using a variety of color data learned from historical documents and literature.
[0332] 4. Reconstructing the frame
[0333] Once the server has captured the colorized frames, it reconstructs them into their original video format, resulting in a colorized video file. The reconstruction process also integrates audio data if necessary.
[0334] 5. Providing a download link
[0335] The server stores the reconstructed color video file and generates a corresponding download link, which is provided to the user, who uses it to download the colorized video file to their device.
[0336] Specific examples
[0337] For example, consider a scenario in which a user wants to upload a monochrome video entitled "Life in Ancient Rome." The user accesses the server's upload page through a web browser on their own device and uploads the video file.
[0338] The server receives the "Life in Ancient Rome" video, splits it into frames, standardizes the resolution of each frame to 720p, removes noise, and adjusts the contrast, preparing the pre-processed frames.
[0339] The frames are then fed into a generative AI model on the server, which then colors each frame with the colors most appropriate for it, based on historical documents about the colors of Roman architecture and clothing, for example.
[0340] The colorized frames are then reassembled and saved as a single video file. Finally, users can use the download link provided by the server to download the colorized version of "Life in Ancient Rome" to their device and watch it.
[0341] This system enables the realistic colorization of black-and-white footage in a short amount of time, and is expected to significantly enhance the visual value of historical film works.
[0342] The processing flow will be explained below.
[0343] Step 1:
[0344] Users upload monochrome video files to the server from their own devices through a designated interface. This operation is performed using a web browser, accessing the upload page, selecting a local video file from the file selection dialog, and pressing the upload button.
[0345] Step 2:
[0346] The server receives the uploaded video file, splits the video data into frames, and uses a video processing library to extract each frame and temporarily store it in memory.
[0347] Step 3:
[0348] The server performs pre-processing on each extracted frame. First, it resizes it to the specified resolution (e.g., 720p) to ensure uniformity. Then it applies a noise reduction filter to improve the image quality. It also performs contrast adjustments to ensure visual consistency across the frames.
[0349] Step 4:
[0350] The server feeds the pre-processed frames into a trained generative AI model, which runs an algorithm to apply the appropriate color to each frame. The model recognizes certain features and, based on those, selects the appropriate color to apply to the frame.
[0351] Step 5:
[0352] The server then reconstructs the colorized frames into a single video file by using a video processing library to sequence the frames, combine them with the original audio data, and save the resulting video file.
[0353] Step 6:
[0354] The server saves the completed colorized video file in the server, generates a link so that the user can download it, and notifies the user of the generated link.
[0355] Step 7:
[0356] Users can download the colorized video file to their device using the download link provided by the server. When they click the link, the file is automatically saved in the download folder.
[0357] As a concrete example, if a user uploads a black-and-white video titled "Life in Ancient Rome," the video will be colorized through the steps described above. The user can then download the colorized version of "Life in Ancient Rome" and view it on their device. This process makes it possible to provide high-quality colorized video in a short amount of time.
[0358] Example 1
[0359] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0360] Conventional techniques for colorizing black-and-white footage require significant time and effort, requiring manual work by people with specialized skills and knowledge. Furthermore, the quality of the colorization is often inconsistent, resulting in unnatural coloring. This makes it difficult to unlock the value of historical and vintage films. Therefore, there is a need for an efficient system for automatically and realistically colorizing black-and-white footage.
[0361] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0362] In this invention, the server includes a means for uploading video files, a means for extracting and preprocessing frames from the video files, a means for colorizing the preprocessed frames using a generative AI model, a means for reconstructing the colorized frames and saving them as video files, and a means for providing users with a download link for the saved video files, thereby enabling realistic colorization of monochrome videos in a short amount of time.
[0363] "Video file" refers to video data stored in digital format.
[0364] "Server" refers to a computer system that provides services over a network.
[0365] "Uploading" refers to the act of transferring data from a local device to a server.
[0366] "Frame" refers to the individual still images that make up a video file.
[0367] "Preprocessing" refers to the preliminary preparation work carried out to improve data quality and for analysis.
[0368] A "generative AI model" refers to a program created using machine learning algorithms that learns patterns from data and generates new data.
[0369] "Colorization" refers to the process of adding color to black and white video or images.
[0370] "Reconstructing" refers to reorganizing the processed data into a new format.
[0371] "Storing" refers to recording and maintaining digital data in a storage device.
[0372] "Download link" refers to a URL for obtaining a digital file over a network.
[0373] "Unifying the resolution" refers to the process of unifying the number of pixels in an image or video according to a certain standard.
[0374] "Noise removal" refers to the process of removing unnecessary random fluctuations from image or video data.
[0375] "Contrast adjustment" refers to the process of adjusting the difference between light and dark to improve the visibility of images and videos.
[0376] "Pre-trained data" refers to an existing dataset that a machine learning model has used during its training process.
[0377] "Coloring" refers to adding appropriate color information to a black and white frame.
[0378] "Historical materials" refers to records and documents about past events and cultures.
[0379] "Realistic colors" refers to realistic and natural color representation.
[0380] This invention is a system for automatically colorizing monochrome movies and video works. This system operates by linking the user's device, a server, and a generative AI model. Specifically, it has the following steps and functions:
[0381] First, the user accesses the server's upload page through a web browser on their device. There, they select a monochrome video file stored on their local disk and upload it to the server. For example, they use the web browser's "file selection" dialog. The file selected by the user is sent to the server via an HTTP POST request.
[0382] When the server receives the uploaded video file, it uses the video processing library FFmpeg to split the video file into frames. Each frame undergoes preprocessing and then the following steps:
[0383] 1. Resolution unification: Resize the resolution of each frame to 720p.
[0384] 2. Noise Reduction: Use a Gaussian filter to remove noise in the video.
[0385] 3. Contrast Adjustment: Perform histogram equalization and adjust the contrast of the frame.
[0386] Once preprocessing is complete, the colorization process is performed using a generative AI model. The server inputs the preprocessed frames into the AI model, which then predicts and applies the optimal color for each frame. The generative AI model used is built using TensorFlow and PyTorch and is trained on a variety of data based on historical documents and literature. For example, based on the training data that "Roman buildings use yellowish stone," a yellowish color is applied to a building.
[0387] The colorized frames are reconstructed again using FFmpeg, and all frames are joined together to form a continuous video file. At this time, if there is audio data from the original video, this is also integrated at the same time. This completes the colorized video file. Specifically, the following FFmpeg command is used to join the frames:
[0388] ffmpeg -r 24 -i frame%04d_colored.png -i original_audio.mp3 -c:v libx264 -c:a aac output_colored_video.mp4
[0389] Finally, the server stores the reconstructed color video file and generates a downloadable link that can be provided to the user, who can use it to download the colorized video file to their device and view it.
[0390] As a concrete example, consider the case where a user uploads a monochrome video of "life in ancient Rome" to a server. The user accesses the server's upload page using a web browser on their device and uploads the "life in ancient Rome" video file. The server then divides the video into frames and preprocesses each frame. The preprocessed frames are then fed into a generative AI model for colorization. For example, each frame is colored with the optimal color based on historical literature on the colors of Roman architecture and clothing. The colorized frames are then reconstructed to generate the final color video file. The user can then download the colorized video file to their device using the provided download link and watch it. This system is expected to rapidly achieve realistic colorization of monochrome video, significantly improving the visual value of historical video works.
[0391] The above is a specific embodiment of the present invention, and the operation and configuration of each part of the system have been specifically described. This system enables realistic and efficient colorization of monochrome footage, significantly reducing the time required compared to traditional manual colorization. Furthermore, the automated process ensures consistently high colorization quality, significantly enhancing the visual appeal of historical footage and vintage films.
[0392] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0393] Step 1:
[0394] The user launches a web browser on their device and accesses the server's upload page. Specifically, they enter the server's upload page URL (e.g., https: / / example.com / upload) in the browser's URL bar. When the page appears, they click the "Choose File" button, select a monochrome video file (e.g., roman_life.mp4) from their local disk, and press the "Upload" button. The input data is a monochrome video file, and the output is a video file uploaded to the server.
[0395] Step 2:
[0396] The server receives the uploaded video file. Then, it splits the video file into frames using FFmpeg. Specifically, it executes the following command:
[0397] ffmpeg -i roman_life.mp4 frame%04d.png
[0398] As a result, the input data is a video file, and the output data is the divided frame images (e.g., frame0001.png, frame0002.png, ...).
[0399] Step 3:
[0400] The server performs pre-processing for each divided frame, which includes the following specific steps:
[0401] Resolution uniformity: To resize each frame to 720p, for example, run the following command:
[0402] ffmpeg -i frame%04d.png -vf scale=1280:720 frame%04d_resized.png
[0403] Denoising: To remove noise using a Gaussian filter, run the following command:
[0404] ffmpeg -i frame%04d_resized.png -vf "gblur=sigma=2" frame%04d_denoised.png
[0405] Contrast adjustment: To perform histogram equalization, run the following command:
[0406] ffmpeg -i frame%04d_denoised.png -vf "eq=contrast=1.5" frame%04d_processed.png
[0407] The input data is a divided frame image, and the output data is a frame image that has been preprocessed.
[0408] Step 4:
[0409] The server inputs the preprocessed frames into a generative AI model for colorization. Specifically, a generative AI model built with TensorFlow or PyTorch is used. The prompt input to the model is in the form of "Please apply appropriate colors based on the frame image I will provide." The input data is the preprocessed frame image, and the output data is the colorized frame image.
[0410] Step 5:
[0411] The server then reconstructs the colorized frames and combines them into a single video file by running the following FFmpeg command:
[0412] ffmpeg -r 24 -i frame%04d_colored.png -i original_audio.mp3 -c:v libx264 -c:a aac output_colored_video.mp4
[0413] The input data are colorized frame images and the original audio data, and the output data is a reconstructed color video file.
[0414] Step 6:
[0415] The server stores the reconstructed color video file and generates a download link, which is provided to the user, who uses it to download the colorized video file to their device. The input data is the reconstructed color video file, and the output data is the download link.
[0416] (Application example 1)
[0417] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0418] Today, there is a growing need to colorize monochrome films and photographs to breathe new life into them. However, traditional colorization methods are time-consuming, costly, and often require specialized knowledge. This creates a need for a system that allows anyone to easily convert monochrome images to color. There is also a need for this functionality to be easily accessible anywhere via mobile devices such as smartphones and tablets.
[0419] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0420] In this invention, the server includes means for uploading video files to the server, means for extracting and preprocessing frames from the video files, means for colorizing the preprocessed frames using a generative AI model, means for reconstructing the colorized frames and saving them as video files, means for providing users with download links for the saved video files, and means for uploading the video files via a smartphone application and making the results available for download. This allows users to easily colorize monochrome videos using their smartphones and quickly obtain the results.
[0421] A "video file" is a digital file containing visual and audio information.
[0422] A "server" is a dedicated computer system that provides services to other computers over a computer network.
[0423] A "frame" is a unit of still images that make up a video file.
[0424] "Preprocessing" refers to processing performed before data analysis or processing, and is an operation performed to improve the quality of the data.
[0425] A "generative AI model" is a trained computational model that uses artificial intelligence technology to analyze data and output results.
[0426] "Colorization" is the process of adding color to monochrome images or videos.
[0427] "Reconstruction" is the process of reassembling decomposed or preprocessed data back into its original form.
[0428] A "download link" is a URL that allows a user to obtain a file via the Internet.
[0429] "User" means a person or entity that uses the system or service.
[0430] A "smartphone application" is a software program that runs on a mobile device.
[0431] This invention is a system for automatically colorizing monochrome video works, and operates in cooperation with a user, a server, and a terminal. Specific embodiments of the invention are described below.
[0432] First, the user launches the application on their smartphone and selects the monochrome video file they want to colorize. This video file is then uploaded to the server via the smartphone application. The uploaded video file is then saved on the server, and the next process begins.
[0433] When the server receives the video file, it divides it into frames and preprocesses each frame. This preprocessing includes unifying the resolution, removing noise, and adjusting the contrast. This is done using software libraries such as OpenCV. Each preprocessed frame is then colorized using a generative AI model. This generative AI model has been trained in advance on a large amount of data, enabling highly accurate colorization.
[0434] The colorized frames are then reconstructed on the server and converted back into the original video file format. The reconstructed color video file is stored on the server, and a download link is generated that users can access. By providing this link to users, they can download the colorized video file via their smartphone.
[0435] Through this process, users can easily colorize monochrome images using their smartphones and receive the results quickly.
[0436] As a concrete example, consider a scenario in which a user uploads a monochrome video titled "Video of an Ancient City." The user first selects and uploads the video file through a smartphone application. The server receives the video file, splits it into frames using OpenCV, standardizes the resolution, removes noise, and adjusts the contrast. Then, a generative AI model colorizes the video, assigning appropriate colors to each frame. The server then reconstructs the colorized frames and generates a download link to provide to the user. Using this link, the user can download and watch the colorized video.
[0437] Another example of a prompt is, "Please input a monochrome frame and generate appropriate colors based on that frame. For example, the building should be the color of stone, the sky should be blue, and the ground should be brown," and the AI model will generate appropriate colors.
[0438] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0439] Step 1:
[0440] A user launches a smartphone application, selects a monochrome video file, and uploads it. This operation causes the smartphone application to send the selected video file to the server. The input is the user's video file, and the output is the file uploaded to the server. At this time, an HTTP POST request is used to send the file to the server.
[0441] Step 2:
[0442] The server receives the uploaded video file and splits it into frames. The input is the video file uploaded to the server, and the output is the split frames. The server uses the OpenCV library to perform the specific operation of splitting the video file into frames.
[0443] Step 3:
[0444] The server preprocesses each frame. The input is a set of split frames, and the output is a set of preprocessed frames. Preprocessing includes unifying the resolution, removing noise, and adjusting the contrast. Specifically, OpenCV is used to unify the resolution of each frame to 720p, apply a noise removal filter, and adjust the contrast using histogram equalization.
[0445] Step 4:
[0446] The server inputs the preprocessed frames into a generative AI model for colorization. The input is a set of preprocessed frames, and the output is a set of colorized frames. The generative AI model applies the appropriate color to each frame based on pre-trained data. Specifically, it uses TensorFlow or PyTorch models to colorize each frame.
[0447] Step 5:
[0448] The server reconstructs the colorized frames and returns them to the original video file format. The input is a set of colorized frames, and the output is a reconstructed color video file. The server uses libraries such as OpenCV to perform the specific operations of successively combining each frame and reconstructing it into the original video format.
[0449] Step 6:
[0450] The server saves the reconstructed color video file and generates a download link for users to access. The input is the reconstructed color video file, and the output is the download link. The server then performs the specific operation of generating a URL for the saved video file and providing it to the user.
[0451] Step 7:
[0452] A user accesses a download link through a smartphone application and downloads a colorized video file. The input is the download link, and the output is the colorized video file on the smartphone. The user clicks on the generated link and performs the specific action of downloading the file using an HTTP GET request.
[0453] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0454] The present invention relates to a system that automatically colorizes monochrome movies and video works by combining an emotion engine with a system that adjusts colors and filters based on the user's emotions. Specific embodiments of the present invention will be described below.
[0455] System Overview
[0456] 1. Uploading video files
[0457] The user uploads monochrome video files to the server through a specified interface from their own terminal. The user accesses the server's upload page using a web browser, selects a local video file using a file selection dialog, and performs the upload operation.
[0458] 2. Frame extraction and preprocessing
[0459] When the server receives the uploaded video file, it divides the video data into frames. Each frame undergoes preprocessing to unify the resolution, remove noise, and adjust the contrast, making it suitable for subsequent AI processing.
[0460] 3. Analysis by Emotion Engine
[0461] An emotion engine installed on the server recognizes the user's emotions. The emotion engine collects the user's facial expression data, voice data, or biometric information, and analyzes this data to identify the user's emotions. For example, if the user shows a happy expression while watching, that emotion data is analyzed.
[0462] 4. Adjusting the colorization process
[0463] Based on the analysis results of the emotion engine, the colorization parameters for the preprocessed frames are adjusted. For example, if the user is emotional, warm colors are emphasized. The preprocessed frames are input into the generative AI model, which then colorizes them based on the adjusted parameters.
[0464] 5. Reconstructing the frame
[0465] The server then reassembles the colorized frames into a single video file, resulting in a complete colorized video file. The reconstruction process also integrates audio data, if necessary.
[0466] 6. Providing download links
[0467] The server stores the reconstructed color video file in the server and generates a link for the user to download it. The created link is notified to the user, who uses it to download the colorized video file to their terminal.
[0468] Specific examples
[0469] For example, consider a scenario in which a user uploads a black and white video entitled "Life in Ancient Rome." The user accesses the server's upload page through a web browser and uploads this video file.
[0470] The server receives the "Life in Ancient Rome" video and splits it into frames, then standardizes the resolution of each frame to 720p, denoises it, and adjusts the contrast, preparing the pre-processed frames.
[0471] If the emotion engine analyzes the user's facial expressions and voice and recognizes that the user has positive emotions toward the video, it adjusts the colorization parameters based on this emotion data to emphasize warm colors.
[0472] The pre-processed frames are then fed into a generative AI model, which then colorizes them based on the adjusted parameters. For example, each frame is colored with the optimal color based on historical documents about the colors of Roman architecture and clothing.
[0473] The colorized frames are then reassembled and saved as a single video file. Finally, users can download the colorized version of "Life in Ancient Rome" to their devices using a download link provided by the server.
[0474] The present invention is expected to provide realistic colorized images based on the user's emotions in a short time, significantly improving visual value and user experience.
[0475] The processing flow will be explained below.
[0476] Step 1:
[0477] Users upload monochrome video files to the server from their own devices through a designated interface. This operation is performed using a web browser, accessing the upload page, selecting a local video file from the file selection dialog, and pressing the upload button.
[0478] Step 2:
[0479] The server receives the uploaded video file and temporarily stores it in a specified directory on the server, before moving on to the next processing step.
[0480] Step 3:
[0481] The server splits the video file into frames, analyzes the video frame by frame using a video processing library, extracts each frame as an image file, and stores it in temporary storage.
[0482] Step 4:
[0483] The server performs pre-processing on each extracted frame. First, it resizes it to the specified resolution (e.g., 720p) to ensure uniformity. Then it applies a noise reduction filter to improve the image quality. It also performs contrast adjustments to ensure visual consistency across the frames.
[0484] Step 5:
[0485] The emotion engine installed on the server recognizes the user's emotions. It collects facial expressions and voice data from the video the user is watching, and analyzes this data to identify the user's emotions. The emotion engine uses facial recognition and voice analysis technology to determine the user's emotional state (happiness, surprise, sadness, etc.).
[0486] Step 6:
[0487] The server automatically adjusts the colorization parameters based on the analysis results of the emotion engine. For example, if the user is recognized as happy, the colors are set to emphasize warm colors. The parameters provided by the emotion engine become input data for the generative AI model.
[0488] Step 7:
[0489] The server inputs the preprocessed frames into the generative AI model based on the adjusted parameters. The AI model selects the optimal color for each frame and colorizes it. The model then references the trained dataset and runs an algorithm to generate realistic colors.
[0490] Step 8:
[0491] The server reconstructs the colorized frames into a single video file, sequences the colorized frames using a video processing library, combines them with the original audio data, and saves the resulting video file.
[0492] Step 9:
[0493] The server stores the generated color image file in the server and generates a link for the user to download it, and notifies the user of the generated link.
[0494] Step 10:
[0495] Users can download the colorized video file to their device using the download link provided by the server. When they click the link, the file is automatically saved in the download folder.
[0496] Specific examples
[0497] For example, if a user uploads a monochrome video titled "Life in Ancient Rome," the video will be colorized through the steps described above. If the emotion engine analyzes the user's facial expressions and voice data and determines that the user has a positive emotion toward the video, the colorization will be adjusted to emphasize warm colors. This process allows the user to download and watch "Life in Ancient Rome," with enhanced visual value and emotional empathy. The present invention is expected to provide realistic colorized video tailored to the emotions of individual users in a short time, significantly improving visual value and user experience.
[0498] Example 2
[0499] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0500] Conventional colorization technologies simply colorize monochrome images and are unable to reflect the user's emotions or visual preferences. As a result, there is a gap between the colorized image and the user's emotions, resulting in a degradation of the viewing experience. Furthermore, due to inconsistent colorization accuracy and quality, parts of the image can appear unnatural.
[0501] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0502] In this invention, the server includes means for uploading a video file to an information processing device, means for extracting and preprocessing frames from the video file, means for collecting and analyzing user emotional information, means for adjusting colorization parameters for the preprocessed frames based on the analysis results, means for colorizing the preprocessed frames using a generative AI model, means for reconstructing the colorized frames and saving them as video files, and means for providing the user with a download link for the saved video file.
[0503] This allows for the provision of more realistic and moving colorized images based on the user's emotions, improving the quality of the viewing experience.
[0504] An "information processing device" is a device for collecting, analyzing, storing, and transmitting data, and includes a server or computer system.
[0505] "Video file" is a digital file containing video data, including monochrome video uploaded by users and final colorized video files.
[0506] A "frame" is a unit of still images that make up a video, and is the basic unit into which a video file is divided when it is processed.
[0507] "Preprocessing" refers to the processing performed on video frames, including resolution unification, noise removal, contrast adjustment, etc.
[0508] "Emotion information" is data that indicates the user's emotional state, and includes facial expression data, voice data, biometric information, and the like.
[0509] An "emotion engine" is a system that analyzes a user's emotional information and identifies their emotional state.
[0510] "Colorization parameters" are setting values for adjusting the color of the frame, and are determined based on the emotion analysis results.
[0511] A "generative AI model" is an artificial intelligence model that colorizes video frames based on pre-trained data.
[0512] "Reconstruction" is the process of connecting the colorized frames in order to recreate the original video file.
[0513] "Download link" refers to a URL or hyperlink that allows a user to download a stored video file via the Internet.
[0514] "Noise reduction" is the process of removing unwanted noise from a video frame.
[0515] "Contrast adjustment" is the process of improving the visibility of an image by adjusting the brightness of the frame.
[0516] "Upload" is the act of a user sending data from their own terminal to a server.
[0517] The present invention relates to a system that automatically colorizes monochrome movies and video works by combining an emotion engine with a system that adjusts colors and filters based on the user's emotions. Specific embodiments of the present invention will be described below.
[0518] (1. Uploading video files)
[0519] The user accesses the upload page of the information processing device using a web browser. The user clicks the "Select File" button on the page and selects a monochrome video file from the local file system. Next, the user presses the upload button to send the video file to the server. The server stores the received video file in a database.
[0520] (2. Frame Extraction and Preprocessing)
[0521] The server passes the received video file to the video analysis module, which divides the video file into frames and sends each frame to the image processing module, which standardizes the resolution to 720p and performs preprocessing such as noise reduction and contrast adjustment. This generates frames suitable for subsequent processing.
[0522] (3. Analysis by Emotion Engine)
[0523] When viewing the video, the user allows the device's camera and microphone to be used. The server collects the user's facial expression data, voice data, and biometric information (e.g., heart rate) in real time. This data is passed to the emotion engine, which analyzes the user's emotions. For example, if the user smiles, the emotion engine analyzes it as "joy."
[0524] (4. Adjusting colorization processing)
[0525] The server's colorization module receives the emotion analysis results from the emotion engine and adjusts the colorization parameters for the preprocessed frames. For example, if the user is emotional, the colorization parameters are set to emphasize warm colors. These adjusted parameters are input as prompts to the generative AI model to colorize each frame. An example prompt is: "The user is emotional. Please use warmer colors."
[0526] (5. Reconstructing the frame)
[0527] The server then reconstructs the colorized frames into a single video file. The reconstruction module then concatenates the frames together to form a continuous video file. If necessary, audio data is also integrated at this stage.
[0528] (6. Providing a download link)
[0529] The server saves the reconstructed color video file in its own storage. It then automatically generates a link that allows the user to download the video file. The user is notified of this link, and can click it to download the colorized video file to their device.
[0530] Specific examples
[0531] For example, consider a scenario in which a user wants to upload a black-and-white video entitled "Life in Ancient Rome." The user uses a web browser to access the upload page of the information processing device and uploads this video file.
[0532] The server receives the "Life in Ancient Rome" video and splits it into frames. Then it standardizes the resolution of each frame to 720p, removes noise, and adjusts the contrast. The pre-processed frames are prepared.
[0533] If the emotion engine analyzes the user's facial expressions and voice and determines that the user has positive emotions toward the video, it adjusts the colorization parameters based on this emotion data to emphasize warm colors. Example prompt: "The user has positive emotions. Use warm colors."
[0534] The pre-processed frames are then fed into a generative AI model, which then colorizes them based on the adjusted parameters. For example, each frame is colored with the optimal color based on historical documents about the colors of Roman architecture and clothing.
[0535] The colorized frames are then reassembled and saved as a single video file. Finally, users can download the colorized version of "Life in Ancient Rome" to their devices using a download link provided by the server.
[0536] The present invention is expected to provide realistic colorized images based on the user's emotions, significantly improving visual satisfaction and user experience.
[0537] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0538] Step 1:
[0539] The user launches a web browser and accesses the upload page of the information processing device. The user clicks the "Select File" button and selects a monochrome video file from the local file system. The user then presses the upload button to send the video file to the server. The input is the video file on the user's local system, and the output is the video file saved on the server.
[0540] Step 2:
[0541] The server passes the received video file to the video analysis module. The video analysis module divides the video file into frames. Each frame is then sent to the image processing module. The image processing module standardizes the resolution to 720p and performs preprocessing such as noise reduction and contrast adjustment. The input is the uploaded video file, and the output is a set of preprocessed frames.
[0542] Step 3:
[0543] When viewing a video, the user allows the device's camera and microphone to be used. The server collects the user's facial expression data, voice data, and biometric information (e.g., heart rate) in real time. This data is passed to the emotion engine. The emotion engine analyzes the user's emotions and generates the analysis results. The input is the user's real-time biometric data, and the output is the user's emotion analysis results.
[0544] Step 4:
[0545] The server's colorization module receives the emotion analysis results from the emotion engine and adjusts the colorization parameters for the preprocessed frames. The emotion analysis results are provided as input, and a prompt sentence is generated based on the results. This prompt sentence is input into a generative AI model to colorize each frame. For example, if the user is emotional, the colorization parameters are set to emphasize warm colors. The input is the emotion analysis results, and the output is the prompt sentence with the adjusted colorization parameters.
[0546] Step 5:
[0547] The server uses a generative AI model to colorize the preprocessed frames using adjusted parameters. Specifically, the generative AI model adds color to each frame based on the prompt. For example, each frame is assigned an appropriate color based on historical documents about the colors of Roman architecture and clothing. The input is the preprocessed frames and the prompt, and the output is a set of colorized frames.
[0548] Step 6:
[0549] The server then reconstructs the colorized frames into a single video file. The reconstruction module concatenates the frames together to form a continuous video file. If necessary, audio data is also integrated at this stage. The input is a set of colorized frames, and the output is a reconstructed color video file.
[0550] Step 7:
[0551] The server saves the reconstructed color video file in its own storage. It then automatically generates a link that allows the user to download the video file. This link is notified to the user, who clicks the link to download the colorized video file to their device. The input is the reconstructed color video file, and the output is the download link.
[0552] (Application example 2)
[0553] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0554] Conventional systems for colorizing black-and-white movies and video works simply add color, and do not provide a personalized viewing experience based on the user's emotions. As a result, there is a problem that the viewing experience is limited because the color adjustment does not correspond to the viewer's emotions. To solve this problem, a system that can colorize in accordance with the user's emotions is needed.
[0555] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading a video file to the information processing device, means for extracting and preprocessing frames from the video file, means for colorizing the preprocessed frames using a generative AI model, means for analyzing a user's emotions, means for adjusting colorization parameters based on the analysis results, means for reconstructing the colorized frames and saving them as a video file, and means for providing a download link for the saved video file. This enables real-time colorization based on the user's emotions.
[0556] "Video file" means a data file that stores the content of visual media in digital format.
[0557] An "information processing device" is a device that processes and manages data, and includes computer systems such as servers.
[0558] A "frame" is a unit of each still image that makes up a video.
[0559] "Preprocessing" refers to the process of preparing data for subsequent processing.
[0560] A "generative AI model" is a model trained using artificial intelligence to perform a specific task, in this case colorization.
[0561] "Means for analyzing emotions" refers to a processing method for identifying emotions based on the user's facial expressions, voice, biometric information, etc.
[0562] A "parameter" is a setting that adjusts the behavior of a system or model.
[0563] "Reconstruction" means reassembling divided data into a single piece of data.
[0564] "Download link" refers to the access means by which a user can save the required data on their device via the Internet.
[0565] The present invention relates to a system for automatically colorizing monochrome movies and video works, and adjusting colors and filters based on the user's emotions. This system is configured and operates as follows.
[0566] First, a user uploads a monochrome video file to the information processing device through a specified interface on their own terminal. The user accesses the upload page using a web browser, selects a local video file using a file selection dialog, and performs the upload operation.
[0567] Next, when the server receives the uploaded video file, it divides the video data into frames. Each frame undergoes preprocessing to unify the resolution, remove noise, and adjust the contrast, making it suitable for subsequent AI processing.
[0568] The preprocessed frames are then analyzed by an emotion analysis engine installed on the server to determine the user's emotional state. The emotion analysis uses the user's facial expression data, voice data, or biometric information. For example, the emotion analysis engine uses cloud services such as EmotionAPI or Amazon Rekognition.
[0569] Based on the analysis, the server generates colorization parameters for the preprocessed frames. A generative AI model (such as OpenAI's GPT-4 or DALL-E) is used to colorize each frame based on these parameters. For example, if the user shows a happy expression, warm colors will be emphasized.
[0570] The colorized frames are then reconstructed into a single video file. During the reconstruction process, audio data is also integrated if necessary. Finally, the colorized video file is stored on a server, and a link is generated so that the user can download it. The created download link is notified to the user, who can use it to download the colorized video file to their device and view it.
[0571] Below are some example prompts that can be used as input to a generative AI model:
[0572] prompt:
[0573] Colorizing black and white "Classic Movies" while analyzing the user's emotional state:
[0574] If the user's emotion is "joy," use warm colors (e.g., warm orange tones)
[0575] If the user is emotional, use soft pastel colors.
[0576] Match the colors of each scene to their historical equivalents
[0577] Detailed frame-by-frame colorization
[0578] This makes it possible to provide an optimal video experience that suits the user's emotions.
[0579] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0580] Step 1:
[0581] A user uploads a monochrome video file to the server through a web browser. The user accesses a web page, selects a video file from a local file selection dialog, and uploads it. The input of this operation is a local monochrome video file, and the output is a video file saved on the server.
[0582] Step 2:
[0583] The server splits the uploaded video file into frames. The server analyzes the video file and extracts each frame. The input is the video file, and the output is a collection of individual frames.
[0584] Step 3:
[0585] The server pre-processes each frame, unifying the resolution of each frame, denoising, and adjusting the contrast. The input is a set of frames, and the output is a set of pre-processed frames.
[0586] Step 4:
[0587] The server analyzes the user's emotions from the preprocessed frames. Using an emotion analysis engine, the server collects and analyzes the user's facial expression data, voice data, and biometric information. The input is the user's facial expression data, voice data, and biometric information, and the output is the user's emotional state.
[0588] Step 5:
[0589] The server generates colorization parameters based on the analysis results. Based on the generated emotion data, the generative AI model generates a prompt sentence, and sets the colorization parameters for each frame based on that prompt sentence. The input is the user's emotional state, and the output is the colorization parameters.
[0590] Step 6:
[0591] The server inputs the preprocessed frames into a generative AI model to perform colorization. The generative AI model (e.g., GPT-4 or DALL-E) colorizes each frame based on the adjusted parameters. The input is a set of preprocessed frames and colorization parameters, and the output is a set of colorized frames.
[0592] Step 7:
[0593] The server reconstructs the colorized frames and saves them as the final color video file. It performs inter-frame reconstruction and integrates audio data as needed. The input is a set of colorized frames and the audio data from the original video, and the output is a reconstructed color video file.
[0594] Step 8:
[0595] The server generates a download link for the reconstructed color video file and provides it to the user. The download link is sent to the user via email or notification. The input is the reconstructed color video file, and the output is the download link provided to the user.
[0596] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0597] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0598] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0599] [Third embodiment]
[0600] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0601] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0602] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0603] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0604] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0605] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0606] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0607] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0608] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0609] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0610] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0611] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0612] The present invention is a system for automatically colorizing monochrome movies and video works, and has the following procedures and functions: A user, a server, and a terminal work together.
[0613] System Overview
[0614] 1. Uploading video files
[0615] The user uploads monochrome video files to the server through a specified interface from their own terminal. The user accesses the server's upload page using a web browser, selects a local video file using a file selection dialog, and performs the upload operation.
[0616] 2. Frame extraction and preprocessing
[0617] When the server receives the uploaded video file, it first divides it into frames. Each frame undergoes pre-processing to unify the resolution, remove noise, and adjust the contrast, making it suitable for subsequent AI processing.
[0618] 3. Colorization
[0619] The pre-processed frames are input into a generative AI model by the server. The AI model applies appropriate colors to each frame based on pre-trained data. The model achieves realistic colorization using a variety of color data learned from historical documents and literature.
[0620] 4. Reconstructing the frame
[0621] Once the server has captured the colorized frames, it reconstructs them into their original video format, resulting in a colorized video file. The reconstruction process also integrates audio data if necessary.
[0622] 5. Providing a download link
[0623] The server stores the reconstructed color video file and generates a corresponding download link, which is provided to the user, who uses it to download the colorized video file to their device.
[0624] Specific examples
[0625] For example, consider a scenario in which a user wants to upload a monochrome video entitled "Life in Ancient Rome." The user accesses the server's upload page through a web browser on their own device and uploads the video file.
[0626] The server receives the "Life in Ancient Rome" video, splits it into frames, standardizes the resolution of each frame to 720p, removes noise, and adjusts the contrast, preparing the pre-processed frames.
[0627] The frames are then fed into a generative AI model on the server, which then colors each frame with the colors most appropriate for it, based on historical documents about the colors of Roman architecture and clothing, for example.
[0628] The colorized frames are then reassembled and saved as a single video file. Finally, users can use the download link provided by the server to download the colorized version of "Life in Ancient Rome" to their device and watch it.
[0629] This system enables the realistic colorization of black-and-white footage in a short amount of time, and is expected to significantly enhance the visual value of historical film works.
[0630] The processing flow will be explained below.
[0631] Step 1:
[0632] Users upload monochrome video files to the server from their own devices through a designated interface. This operation is performed using a web browser, accessing the upload page, selecting a local video file from the file selection dialog, and pressing the upload button.
[0633] Step 2:
[0634] The server receives the uploaded video file, splits the video data into frames, and uses a video processing library to extract each frame and temporarily store it in memory.
[0635] Step 3:
[0636] The server performs pre-processing on each extracted frame. First, it resizes it to the specified resolution (e.g., 720p) to ensure uniformity. Then it applies a noise reduction filter to improve the image quality. It also performs contrast adjustments to ensure visual consistency across the frames.
[0637] Step 4:
[0638] The server feeds the pre-processed frames into a trained generative AI model, which runs an algorithm to apply the appropriate color to each frame. The model recognizes certain features and, based on those, selects the appropriate color to apply to the frame.
[0639] Step 5:
[0640] The server then reconstructs the colorized frames into a single video file by using a video processing library to sequence the frames, combine them with the original audio data, and save the resulting video file.
[0641] Step 6:
[0642] The server saves the completed colorized video file in the server, generates a link so that the user can download it, and notifies the user of the generated link.
[0643] Step 7:
[0644] Users can download the colorized video file to their device using the download link provided by the server. When they click the link, the file is automatically saved in the download folder.
[0645] As a concrete example, if a user uploads a black-and-white video titled "Life in Ancient Rome," the video will be colorized through the steps described above. The user can then download the colorized version of "Life in Ancient Rome" and view it on their device. This process makes it possible to provide high-quality colorized video in a short amount of time.
[0646] Example 1
[0647] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0648] Conventional techniques for colorizing black-and-white footage require significant time and effort, requiring manual work by people with specialized skills and knowledge. Furthermore, the quality of the colorization is often inconsistent, resulting in unnatural coloring. This makes it difficult to unlock the value of historical and vintage films. Therefore, there is a need for an efficient system for automatically and realistically colorizing black-and-white footage.
[0649] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0650] In this invention, the server includes a means for uploading video files, a means for extracting and preprocessing frames from the video files, a means for colorizing the preprocessed frames using a generative AI model, a means for reconstructing the colorized frames and saving them as video files, and a means for providing users with a download link for the saved video files, thereby enabling realistic colorization of monochrome videos in a short amount of time.
[0651] "Video file" refers to video data stored in digital format.
[0652] "Server" refers to a computer system that provides services over a network.
[0653] "Uploading" refers to the act of transferring data from a local device to a server.
[0654] "Frame" refers to the individual still images that make up a video file.
[0655] "Preprocessing" refers to the preliminary preparation work carried out to improve data quality and for analysis.
[0656] A "generative AI model" refers to a program created using machine learning algorithms that learns patterns from data and generates new data.
[0657] "Colorization" refers to the process of adding color to black and white video or images.
[0658] "Reconstructing" refers to reorganizing the processed data into a new format.
[0659] "Storing" refers to recording and maintaining digital data in a storage device.
[0660] "Download link" refers to a URL for obtaining a digital file over a network.
[0661] "Unifying the resolution" refers to the process of unifying the number of pixels in an image or video according to a certain standard.
[0662] "Noise removal" refers to the process of removing unnecessary random fluctuations from image or video data.
[0663] "Contrast adjustment" refers to the process of adjusting the difference between light and dark to improve the visibility of images and videos.
[0664] "Pre-trained data" refers to an existing dataset that a machine learning model has used during its training process.
[0665] "Coloring" refers to adding appropriate color information to a black and white frame.
[0666] "Historical materials" refers to records and documents about past events and cultures.
[0667] "Realistic colors" refers to realistic and natural color representation.
[0668] This invention is a system for automatically colorizing monochrome movies and video works. This system operates by linking the user's device, a server, and a generative AI model. Specifically, it has the following steps and functions:
[0669] First, the user accesses the server's upload page through a web browser on their device. There, they select a monochrome video file stored on their local disk and upload it to the server. For example, they use the web browser's "file selection" dialog. The file selected by the user is sent to the server via an HTTP POST request.
[0670] When the server receives the uploaded video file, it uses the video processing library FFmpeg to split the video file into frames. Each frame undergoes preprocessing and then the following steps:
[0671] 1. Resolution unification: Resize the resolution of each frame to 720p.
[0672] 2. Noise Reduction: Use a Gaussian filter to remove noise in the video.
[0673] 3. Contrast Adjustment: Perform histogram equalization and adjust the contrast of the frame.
[0674] Once preprocessing is complete, the colorization process is performed using a generative AI model. The server inputs the preprocessed frames into the AI model, which then predicts and applies the optimal color for each frame. The generative AI model used is built using TensorFlow and PyTorch and is trained on a variety of data based on historical documents and literature. For example, based on the training data that "Roman buildings use yellowish stone," a yellowish color is applied to a building.
[0675] The colorized frames are reconstructed again using FFmpeg, and all frames are joined together to form a continuous video file. At this time, if there is audio data from the original video, this is also integrated at the same time. This completes the colorized video file. Specifically, the following FFmpeg command is used to join the frames:
[0676] ffmpeg -r 24 -i frame%04d_colored.png -i original_audio.mp3 -c:v libx264 -c:a aac output_colored_video.mp4
[0677] Finally, the server stores the reconstructed color video file and generates a downloadable link that can be provided to the user, who can use it to download the colorized video file to their device and view it.
[0678] As a concrete example, consider the case where a user uploads a monochrome video of "life in ancient Rome" to a server. The user accesses the server's upload page using a web browser on their device and uploads the "life in ancient Rome" video file. The server then divides the video into frames and preprocesses each frame. The preprocessed frames are then fed into a generative AI model for colorization. For example, each frame is colored with the optimal color based on historical literature on the colors of Roman architecture and clothing. The colorized frames are then reconstructed to generate the final color video file. The user can then download the colorized video file to their device using the provided download link and watch it. This system is expected to rapidly achieve realistic colorization of monochrome video, significantly improving the visual value of historical video works.
[0679] The above is a specific embodiment of the present invention, and the operation and configuration of each part of the system have been specifically described. This system enables realistic and efficient colorization of monochrome footage, significantly reducing the time required compared to traditional manual colorization. Furthermore, the automated process ensures consistently high colorization quality, significantly enhancing the visual appeal of historical footage and vintage films.
[0680] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0681] Step 1:
[0682] The user launches a web browser on their device and accesses the server's upload page. Specifically, they enter the server's upload page URL (e.g., https: / / example.com / upload) in the browser's URL bar. When the page appears, they click the "Choose File" button, select a monochrome video file (e.g., roman_life.mp4) from their local disk, and press the "Upload" button. The input data is a monochrome video file, and the output is a video file uploaded to the server.
[0683] Step 2:
[0684] The server receives the uploaded video file. Then, it splits the video file into frames using FFmpeg. Specifically, it executes the following command:
[0685] ffmpeg -i roman_life.mp4 frame%04d.png
[0686] As a result, the input data is a video file, and the output data is the divided frame images (e.g., frame0001.png, frame0002.png, ...).
[0687] Step 3:
[0688] The server performs pre-processing for each divided frame, which includes the following specific steps:
[0689] Resolution uniformity: To resize each frame to 720p, for example, run the following command:
[0690] ffmpeg -i frame%04d.png -vf scale=1280:720 frame%04d_resized.png
[0691] Denoising: To remove noise using a Gaussian filter, run the following command:
[0692] ffmpeg -i frame%04d_resized.png -vf "gblur=sigma=2" frame%04d_denoised.png
[0693] Contrast adjustment: To perform histogram equalization, run the following command:
[0694] ffmpeg -i frame%04d_denoised.png -vf "eq=contrast=1.5" frame%04d_processed.png
[0695] The input data is a divided frame image, and the output data is a frame image that has been preprocessed.
[0696] Step 4:
[0697] The server inputs the preprocessed frames into a generative AI model for colorization. Specifically, a generative AI model built with TensorFlow or PyTorch is used. The prompt input to the model is in the form of "Please apply appropriate colors based on the frame image I will provide." The input data is the preprocessed frame image, and the output data is the colorized frame image.
[0698] Step 5:
[0699] The server then reconstructs the colorized frames and combines them into a single video file by running the following FFmpeg command:
[0700] ffmpeg -r 24 -i frame%04d_colored.png -i original_audio.mp3 -c:v libx264 -c:a aac output_colored_video.mp4
[0701] The input data are colorized frame images and the original audio data, and the output data is a reconstructed color video file.
[0702] Step 6:
[0703] The server stores the reconstructed color video file and generates a download link, which is provided to the user, who uses it to download the colorized video file to their device. The input data is the reconstructed color video file, and the output data is the download link.
[0704] (Application example 1)
[0705] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0706] Today, there is a growing need to colorize monochrome films and photographs to breathe new life into them. However, traditional colorization methods are time-consuming, costly, and often require specialized knowledge. This creates a need for a system that allows anyone to easily convert monochrome images to color. There is also a need for this functionality to be easily accessible anywhere via mobile devices such as smartphones and tablets.
[0707] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0708] In this invention, the server includes means for uploading video files to the server, means for extracting and preprocessing frames from the video files, means for colorizing the preprocessed frames using a generative AI model, means for reconstructing the colorized frames and saving them as video files, means for providing users with download links for the saved video files, and means for uploading the video files via a smartphone application and making the results available for download. This allows users to easily colorize monochrome videos using their smartphones and quickly obtain the results.
[0709] A "video file" is a digital file containing visual and audio information.
[0710] A "server" is a dedicated computer system that provides services to other computers over a computer network.
[0711] A "frame" is a unit of still images that make up a video file.
[0712] "Preprocessing" refers to processing performed before data analysis or processing, and is an operation performed to improve the quality of the data.
[0713] A "generative AI model" is a trained computational model that uses artificial intelligence technology to analyze data and output results.
[0714] "Colorization" is the process of adding color to monochrome images or videos.
[0715] "Reconstruction" is the process of reassembling decomposed or preprocessed data back into its original form.
[0716] A "download link" is a URL that allows a user to obtain a file via the Internet.
[0717] "User" means a person or entity that uses the system or service.
[0718] A "smartphone application" is a software program that runs on a mobile device.
[0719] This invention is a system for automatically colorizing monochrome video works, and operates in cooperation with a user, a server, and a terminal. Specific embodiments of the invention are described below.
[0720] First, the user launches the application on their smartphone and selects the monochrome video file they want to colorize. This video file is then uploaded to the server via the smartphone application. The uploaded video file is then saved on the server, and the next process begins.
[0721] When the server receives the video file, it divides it into frames and preprocesses each frame. This preprocessing includes unifying the resolution, removing noise, and adjusting the contrast. This is done using software libraries such as OpenCV. Each preprocessed frame is then colorized using a generative AI model. This generative AI model has been trained in advance on a large amount of data, enabling highly accurate colorization.
[0722] The colorized frames are then reconstructed on the server and converted back into the original video file format. The reconstructed color video file is stored on the server, and a download link is generated that users can access. By providing this link to users, they can download the colorized video file via their smartphone.
[0723] Through this process, users can easily colorize monochrome images using their smartphones and receive the results quickly.
[0724] As a concrete example, consider a scenario in which a user uploads a monochrome video titled "Video of an Ancient City." The user first selects and uploads the video file through a smartphone application. The server receives the video file, splits it into frames using OpenCV, standardizes the resolution, removes noise, and adjusts the contrast. Then, a generative AI model colorizes the video, assigning appropriate colors to each frame. The server then reconstructs the colorized frames and generates a download link to provide to the user. Using this link, the user can download and watch the colorized video.
[0725] Another example of a prompt is, "Please input a monochrome frame and generate appropriate colors based on that frame. For example, the building should be the color of stone, the sky should be blue, and the ground should be brown," and the AI model will generate appropriate colors.
[0726] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0727] Step 1:
[0728] A user launches a smartphone application, selects a monochrome video file, and uploads it. This operation causes the smartphone application to send the selected video file to the server. The input is the user's video file, and the output is the file uploaded to the server. At this time, an HTTP POST request is used to send the file to the server.
[0729] Step 2:
[0730] The server receives the uploaded video file and splits it into frames. The input is the video file uploaded to the server, and the output is the split frames. The server uses the OpenCV library to perform the specific operation of splitting the video file into frames.
[0731] Step 3:
[0732] The server preprocesses each frame. The input is a set of split frames, and the output is a set of preprocessed frames. Preprocessing includes unifying the resolution, removing noise, and adjusting the contrast. Specifically, OpenCV is used to unify the resolution of each frame to 720p, apply a noise removal filter, and adjust the contrast using histogram equalization.
[0733] Step 4:
[0734] The server inputs the preprocessed frames into a generative AI model for colorization. The input is a set of preprocessed frames, and the output is a set of colorized frames. The generative AI model applies the appropriate color to each frame based on pre-trained data. Specifically, it uses TensorFlow or PyTorch models to colorize each frame.
[0735] Step 5:
[0736] The server reconstructs the colorized frames and returns them to the original video file format. The input is a set of colorized frames, and the output is a reconstructed color video file. The server uses libraries such as OpenCV to perform the specific operations of successively combining each frame and reconstructing it into the original video format.
[0737] Step 6:
[0738] The server saves the reconstructed color video file and generates a download link for users to access. The input is the reconstructed color video file, and the output is the download link. The server then performs the specific operation of generating a URL for the saved video file and providing it to the user.
[0739] Step 7:
[0740] A user accesses a download link through a smartphone application and downloads a colorized video file. The input is the download link, and the output is the colorized video file on the smartphone. The user clicks on the generated link and performs the specific action of downloading the file using an HTTP GET request.
[0741] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0742] The present invention relates to a system that automatically colorizes monochrome movies and video works by combining an emotion engine with a system that adjusts colors and filters based on the user's emotions. Specific embodiments of the present invention will be described below.
[0743] System Overview
[0744] 1. Uploading video files
[0745] The user uploads monochrome video files to the server through a specified interface from their own terminal. The user accesses the server's upload page using a web browser, selects a local video file using a file selection dialog, and performs the upload operation.
[0746] 2. Frame extraction and preprocessing
[0747] When the server receives the uploaded video file, it divides the video data into frames. Each frame undergoes preprocessing to unify the resolution, remove noise, and adjust the contrast, making it suitable for subsequent AI processing.
[0748] 3. Analysis by Emotion Engine
[0749] An emotion engine installed on the server recognizes the user's emotions. The emotion engine collects the user's facial expression data, voice data, or biometric information, and analyzes this data to identify the user's emotions. For example, if the user shows a happy expression while watching, that emotion data is analyzed.
[0750] 4. Adjusting the colorization process
[0751] Based on the analysis results of the emotion engine, the colorization parameters for the preprocessed frames are adjusted. For example, if the user is emotional, warm colors are emphasized. The preprocessed frames are input into the generative AI model, which then colorizes them based on the adjusted parameters.
[0752] 5. Reconstructing the frame
[0753] The server then reassembles the colorized frames into a single video file, resulting in a complete colorized video file. The reconstruction process also integrates audio data, if necessary.
[0754] 6. Providing download links
[0755] The server stores the reconstructed color video file in the server and generates a link for the user to download it. The created link is notified to the user, who uses it to download the colorized video file to their terminal.
[0756] Specific examples
[0757] For example, consider a scenario in which a user uploads a black and white video entitled "Life in Ancient Rome." The user accesses the server's upload page through a web browser and uploads this video file.
[0758] The server receives the "Life in Ancient Rome" video and splits it into frames, then standardizes the resolution of each frame to 720p, denoises it, and adjusts the contrast, preparing the pre-processed frames.
[0759] If the emotion engine analyzes the user's facial expressions and voice and recognizes that the user has positive emotions toward the video, it adjusts the colorization parameters based on this emotion data to emphasize warm colors.
[0760] The pre-processed frames are then fed into a generative AI model, which then colorizes them based on the adjusted parameters. For example, each frame is colored with the optimal color based on historical documents about the colors of Roman architecture and clothing.
[0761] The colorized frames are then reassembled and saved as a single video file. Finally, users can download the colorized version of "Life in Ancient Rome" to their devices using a download link provided by the server.
[0762] The present invention is expected to provide realistic colorized images based on the user's emotions in a short time, significantly improving visual value and user experience.
[0763] The processing flow will be explained below.
[0764] Step 1:
[0765] Users upload monochrome video files to the server from their own devices through a designated interface. This operation is performed using a web browser, accessing the upload page, selecting a local video file from the file selection dialog, and pressing the upload button.
[0766] Step 2:
[0767] The server receives the uploaded video file and temporarily stores it in a specified directory on the server, before moving on to the next processing step.
[0768] Step 3:
[0769] The server splits the video file into frames, analyzes the video frame by frame using a video processing library, extracts each frame as an image file, and stores it in temporary storage.
[0770] Step 4:
[0771] The server performs pre-processing on each extracted frame. First, it resizes it to the specified resolution (e.g., 720p) to ensure uniformity. Then it applies a noise reduction filter to improve the image quality. It also performs contrast adjustments to ensure visual consistency across the frames.
[0772] Step 5:
[0773] The emotion engine installed on the server recognizes the user's emotions. It collects facial expressions and voice data from the video the user is watching, and analyzes this data to identify the user's emotions. The emotion engine uses facial recognition and voice analysis technology to determine the user's emotional state (happiness, surprise, sadness, etc.).
[0774] Step 6:
[0775] The server automatically adjusts the colorization parameters based on the analysis results of the emotion engine. For example, if the user is recognized as happy, the colors are set to emphasize warm colors. The parameters provided by the emotion engine become input data for the generative AI model.
[0776] Step 7:
[0777] The server inputs the preprocessed frames into the generative AI model based on the adjusted parameters. The AI model selects the optimal color for each frame and colorizes it. The model then references the trained dataset and runs an algorithm to generate realistic colors.
[0778] Step 8:
[0779] The server reconstructs the colorized frames into a single video file, sequences the colorized frames using a video processing library, combines them with the original audio data, and saves the resulting video file.
[0780] Step 9:
[0781] The server stores the generated color image file in the server and generates a link for the user to download it, and notifies the user of the generated link.
[0782] Step 10:
[0783] Users can download the colorized video file to their device using the download link provided by the server. When they click the link, the file is automatically saved in the download folder.
[0784] Specific examples
[0785] For example, if a user uploads a monochrome video titled "Life in Ancient Rome," the video will be colorized through the steps described above. If the emotion engine analyzes the user's facial expressions and voice data and determines that the user has a positive emotion toward the video, the colorization will be adjusted to emphasize warm colors. This process allows the user to download and watch "Life in Ancient Rome," with enhanced visual value and emotional empathy. The present invention is expected to provide realistic colorized video tailored to the emotions of individual users in a short time, significantly improving visual value and user experience.
[0786] Example 2
[0787] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0788] Conventional colorization technologies simply colorize monochrome images and are unable to reflect the user's emotions or visual preferences. As a result, there is a gap between the colorized image and the user's emotions, resulting in a degradation of the viewing experience. Furthermore, due to inconsistent colorization accuracy and quality, parts of the image can appear unnatural.
[0789] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0790] In this invention, the server includes means for uploading a video file to an information processing device, means for extracting and preprocessing frames from the video file, means for collecting and analyzing user emotional information, means for adjusting colorization parameters for the preprocessed frames based on the analysis results, means for colorizing the preprocessed frames using a generative AI model, means for reconstructing the colorized frames and saving them as video files, and means for providing the user with a download link for the saved video file.
[0791] This allows for the provision of more realistic and moving colorized images based on the user's emotions, improving the quality of the viewing experience.
[0792] An "information processing device" is a device for collecting, analyzing, storing, and transmitting data, and includes a server or computer system.
[0793] "Video file" is a digital file containing video data, including monochrome video uploaded by users and final colorized video files.
[0794] A "frame" is a unit of still images that make up a video, and is the basic unit into which a video file is divided when it is processed.
[0795] "Preprocessing" refers to the processing performed on video frames, including resolution unification, noise removal, contrast adjustment, etc.
[0796] "Emotion information" is data that indicates the user's emotional state, and includes facial expression data, voice data, biometric information, and the like.
[0797] An "emotion engine" is a system that analyzes a user's emotional information and identifies their emotional state.
[0798] "Colorization parameters" are setting values for adjusting the color of the frame, and are determined based on the emotion analysis results.
[0799] A "generative AI model" is an artificial intelligence model that colorizes video frames based on pre-trained data.
[0800] "Reconstruction" is the process of connecting the colorized frames in order to recreate the original video file.
[0801] "Download link" refers to a URL or hyperlink that allows a user to download a stored video file via the Internet.
[0802] "Noise reduction" is the process of removing unwanted noise from a video frame.
[0803] "Contrast adjustment" is the process of improving the visibility of an image by adjusting the brightness of the frame.
[0804] "Upload" is the act of a user sending data from their own terminal to a server.
[0805] The present invention relates to a system that automatically colorizes monochrome movies and video works by combining an emotion engine with a system that adjusts colors and filters based on the user's emotions. Specific embodiments of the present invention will be described below.
[0806] (1. Uploading video files)
[0807] The user accesses the upload page of the information processing device using a web browser. The user clicks the "Select File" button on the page and selects a monochrome video file from the local file system. Next, the user presses the upload button to send the video file to the server. The server stores the received video file in a database.
[0808] (2. Frame Extraction and Preprocessing)
[0809] The server passes the received video file to the video analysis module, which divides the video file into frames and sends each frame to the image processing module, which standardizes the resolution to 720p and performs preprocessing such as noise reduction and contrast adjustment. This generates frames suitable for subsequent processing.
[0810] (3. Analysis by Emotion Engine)
[0811] When viewing the video, the user allows the device's camera and microphone to be used. The server collects the user's facial expression data, voice data, and biometric information (e.g., heart rate) in real time. This data is passed to the emotion engine, which analyzes the user's emotions. For example, if the user smiles, the emotion engine analyzes it as "joy."
[0812] (4. Adjusting colorization processing)
[0813] The server's colorization module receives the emotion analysis results from the emotion engine and adjusts the colorization parameters for the preprocessed frames. For example, if the user is emotional, the colorization parameters are set to emphasize warm colors. These adjusted parameters are input as prompts to the generative AI model to colorize each frame. An example prompt is: "The user is emotional. Please use warmer colors."
[0814] (5. Reconstructing the frame)
[0815] The server then reconstructs the colorized frames into a single video file. The reconstruction module then concatenates the frames together to form a continuous video file. If necessary, audio data is also integrated at this stage.
[0816] (6. Providing a download link)
[0817] The server saves the reconstructed color video file in its own storage. It then automatically generates a link that allows the user to download the video file. The user is notified of this link, and can click it to download the colorized video file to their device.
[0818] Specific examples
[0819] For example, consider a scenario in which a user wants to upload a black-and-white video entitled "Life in Ancient Rome." The user uses a web browser to access the upload page of the information processing device and uploads this video file.
[0820] The server receives the "Life in Ancient Rome" video and splits it into frames. Then it standardizes the resolution of each frame to 720p, removes noise, and adjusts the contrast. The pre-processed frames are prepared.
[0821] If the emotion engine analyzes the user's facial expressions and voice and determines that the user has positive emotions toward the video, it adjusts the colorization parameters based on this emotion data to emphasize warm colors. Example prompt: "The user has positive emotions. Use warm colors."
[0822] The pre-processed frames are then fed into a generative AI model, which then colorizes them based on the adjusted parameters. For example, each frame is colored with the optimal color based on historical documents about the colors of Roman architecture and clothing.
[0823] The colorized frames are then reassembled and saved as a single video file. Finally, users can download the colorized version of "Life in Ancient Rome" to their devices using a download link provided by the server.
[0824] The present invention is expected to provide realistic colorized images based on the user's emotions, significantly improving visual satisfaction and user experience.
[0825] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0826] Step 1:
[0827] The user launches a web browser and accesses the upload page of the information processing device. The user clicks the "Select File" button and selects a monochrome video file from the local file system. The user then presses the upload button to send the video file to the server. The input is the video file on the user's local system, and the output is the video file saved on the server.
[0828] Step 2:
[0829] The server passes the received video file to the video analysis module. The video analysis module divides the video file into frames. Each frame is then sent to the image processing module. The image processing module standardizes the resolution to 720p and performs preprocessing such as noise reduction and contrast adjustment. The input is the uploaded video file, and the output is a set of preprocessed frames.
[0830] Step 3:
[0831] When viewing a video, the user allows the device's camera and microphone to be used. The server collects the user's facial expression data, voice data, and biometric information (e.g., heart rate) in real time. This data is passed to the emotion engine. The emotion engine analyzes the user's emotions and generates the analysis results. The input is the user's real-time biometric data, and the output is the user's emotion analysis results.
[0832] Step 4:
[0833] The server's colorization module receives the emotion analysis results from the emotion engine and adjusts the colorization parameters for the preprocessed frames. The emotion analysis results are provided as input, and a prompt sentence is generated based on the results. This prompt sentence is input into a generative AI model to colorize each frame. For example, if the user is emotional, the colorization parameters are set to emphasize warm colors. The input is the emotion analysis results, and the output is the prompt sentence with the adjusted colorization parameters.
[0834] Step 5:
[0835] The server uses a generative AI model to colorize the preprocessed frames using adjusted parameters. Specifically, the generative AI model adds color to each frame based on the prompt. For example, each frame is assigned an appropriate color based on historical documents about the colors of Roman architecture and clothing. The input is the preprocessed frames and the prompt, and the output is a set of colorized frames.
[0836] Step 6:
[0837] The server then reconstructs the colorized frames into a single video file. The reconstruction module concatenates the frames together to form a continuous video file. If necessary, audio data is also integrated at this stage. The input is a set of colorized frames, and the output is a reconstructed color video file.
[0838] Step 7:
[0839] The server saves the reconstructed color video file in its own storage. It then automatically generates a link that allows the user to download the video file. This link is notified to the user, who clicks the link to download the colorized video file to their device. The input is the reconstructed color video file, and the output is the download link.
[0840] (Application example 2)
[0841] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0842] Conventional systems for colorizing black-and-white movies and video works simply add color, and do not provide a personalized viewing experience based on the user's emotions. As a result, there is a problem that the viewing experience is limited because the color adjustment does not correspond to the viewer's emotions. To solve this problem, a system that can colorize in accordance with the user's emotions is needed.
[0843] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading a video file to the information processing device, means for extracting and preprocessing frames from the video file, means for colorizing the preprocessed frames using a generative AI model, means for analyzing a user's emotions, means for adjusting colorization parameters based on the analysis results, means for reconstructing the colorized frames and saving them as a video file, and means for providing a download link for the saved video file. This enables real-time colorization based on the user's emotions.
[0844] "Video file" means a data file that stores the content of visual media in digital format.
[0845] An "information processing device" is a device that processes and manages data, and includes computer systems such as servers.
[0846] A "frame" is a unit of each still image that makes up a video.
[0847] "Preprocessing" refers to the process of preparing data for subsequent processing.
[0848] A "generative AI model" is a model trained using artificial intelligence to perform a specific task, in this case colorization.
[0849] "Means for analyzing emotions" refers to a processing method for identifying emotions based on the user's facial expressions, voice, biometric information, etc.
[0850] A "parameter" is a setting that adjusts the behavior of a system or model.
[0851] "Reconstruction" means reassembling divided data into a single piece of data.
[0852] "Download link" refers to the access means by which a user can save the required data on their device via the Internet.
[0853] The present invention relates to a system for automatically colorizing monochrome movies and video works, and adjusting colors and filters based on the user's emotions. This system is configured and operates as follows.
[0854] First, a user uploads a monochrome video file to the information processing device through a specified interface on their own terminal. The user accesses the upload page using a web browser, selects a local video file using a file selection dialog, and performs the upload operation.
[0855] Next, when the server receives the uploaded video file, it divides the video data into frames. Each frame undergoes preprocessing to unify the resolution, remove noise, and adjust the contrast, making it suitable for subsequent AI processing.
[0856] The preprocessed frames are then analyzed by an emotion analysis engine installed on the server to determine the user's emotional state. The emotion analysis uses the user's facial expression data, voice data, or biometric information. For example, the emotion analysis engine uses cloud services such as EmotionAPI or Amazon Rekognition.
[0857] Based on the analysis, the server generates colorization parameters for the preprocessed frames. A generative AI model (such as OpenAI's GPT-4 or DALL-E) is used to colorize each frame based on these parameters. For example, if the user shows a happy expression, warm colors will be emphasized.
[0858] The colorized frames are then reconstructed into a single video file. During the reconstruction process, audio data is also integrated if necessary. Finally, the colorized video file is stored on a server, and a link is generated so that the user can download it. The created download link is notified to the user, who can use it to download the colorized video file to their device and view it.
[0859] Below are some example prompts that can be used as input to a generative AI model:
[0860] prompt:
[0861] Colorizing black and white "Classic Movies" while analyzing the user's emotional state:
[0862] If the user's emotion is "joy," use warm colors (e.g., warm orange tones)
[0863] If the user is emotional, use soft pastel colors.
[0864] Match the colors of each scene to their historical equivalents
[0865] Detailed frame-by-frame colorization
[0866] This makes it possible to provide an optimal video experience that suits the user's emotions.
[0867] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0868] Step 1:
[0869] A user uploads a monochrome video file to the server through a web browser. The user accesses a web page, selects a video file from a local file selection dialog, and uploads it. The input of this operation is a local monochrome video file, and the output is a video file saved on the server.
[0870] Step 2:
[0871] The server splits the uploaded video file into frames. The server analyzes the video file and extracts each frame. The input is the video file, and the output is a collection of individual frames.
[0872] Step 3:
[0873] The server pre-processes each frame, unifying the resolution of each frame, denoising, and adjusting the contrast. The input is a set of frames, and the output is a set of pre-processed frames.
[0874] Step 4:
[0875] The server analyzes the user's emotions from the preprocessed frames. Using an emotion analysis engine, the server collects and analyzes the user's facial expression data, voice data, and biometric information. The input is the user's facial expression data, voice data, and biometric information, and the output is the user's emotional state.
[0876] Step 5:
[0877] The server generates colorization parameters based on the analysis results. Based on the generated emotion data, the generative AI model generates a prompt sentence, and sets the colorization parameters for each frame based on that prompt sentence. The input is the user's emotional state, and the output is the colorization parameters.
[0878] Step 6:
[0879] The server inputs the preprocessed frames into a generative AI model to perform colorization. The generative AI model (e.g., GPT-4 or DALL-E) colorizes each frame based on the adjusted parameters. The input is a set of preprocessed frames and colorization parameters, and the output is a set of colorized frames.
[0880] Step 7:
[0881] The server reconstructs the colorized frames and saves them as the final color video file. It performs inter-frame reconstruction and integrates audio data as needed. The input is a set of colorized frames and the audio data from the original video, and the output is a reconstructed color video file.
[0882] Step 8:
[0883] The server generates a download link for the reconstructed color video file and provides it to the user. The download link is sent to the user via email or notification. The input is the reconstructed color video file, and the output is the download link provided to the user.
[0884] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0885] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0886] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0887] [Fourth embodiment]
[0888] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0889] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0890] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0891] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0892] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0893] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0894] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0895] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0896] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0897] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0898] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0899] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0900] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0901] The present invention is a system for automatically colorizing monochrome movies and video works, and has the following procedures and functions: A user, a server, and a terminal work together.
[0902] System Overview
[0903] 1. Uploading video files
[0904] The user uploads monochrome video files to the server through a specified interface from their own terminal. The user accesses the server's upload page using a web browser, selects a local video file using a file selection dialog, and performs the upload operation.
[0905] 2. Frame extraction and preprocessing
[0906] When the server receives the uploaded video file, it first divides it into frames. Each frame undergoes pre-processing to unify the resolution, remove noise, and adjust the contrast, making it suitable for subsequent AI processing.
[0907] 3. Colorization
[0908] The pre-processed frames are input into a generative AI model by the server. The AI model applies appropriate colors to each frame based on pre-trained data. The model achieves realistic colorization using a variety of color data learned from historical documents and literature.
[0909] 4. Reconstructing the frame
[0910] Once the server has captured the colorized frames, it reconstructs them into their original video format, resulting in a colorized video file. The reconstruction process also integrates audio data if necessary.
[0911] 5. Providing a download link
[0912] The server stores the reconstructed color video file and generates a corresponding download link, which is provided to the user, who uses it to download the colorized video file to their device.
[0913] Specific examples
[0914] For example, consider a scenario in which a user wants to upload a monochrome video entitled "Life in Ancient Rome." The user accesses the server's upload page through a web browser on their own device and uploads the video file.
[0915] The server receives the "Life in Ancient Rome" video, splits it into frames, standardizes the resolution of each frame to 720p, removes noise, and adjusts the contrast, preparing the pre-processed frames.
[0916] The frames are then fed into a generative AI model on the server, which then colors each frame with the colors most appropriate for it, based on historical documents about the colors of Roman architecture and clothing, for example.
[0917] The colorized frames are then reassembled and saved as a single video file. Finally, users can use the download link provided by the server to download the colorized version of "Life in Ancient Rome" to their device and watch it.
[0918] This system enables the realistic colorization of black-and-white footage in a short amount of time, and is expected to significantly enhance the visual value of historical film works.
[0919] The processing flow will be explained below.
[0920] Step 1:
[0921] Users upload monochrome video files to the server from their own devices through a designated interface. This operation is performed using a web browser, accessing the upload page, selecting a local video file from the file selection dialog, and pressing the upload button.
[0922] Step 2:
[0923] The server receives the uploaded video file, splits the video data into frames, and uses a video processing library to extract each frame and temporarily store it in memory.
[0924] Step 3:
[0925] The server performs pre-processing on each extracted frame. First, it resizes it to the specified resolution (e.g., 720p) to ensure uniformity. Then it applies a noise reduction filter to improve the image quality. It also performs contrast adjustments to ensure visual consistency across the frames.
[0926] Step 4:
[0927] The server feeds the pre-processed frames into a trained generative AI model, which runs an algorithm to apply the appropriate color to each frame. The model recognizes certain features and, based on those, selects the appropriate color to apply to the frame.
[0928] Step 5:
[0929] The server then reconstructs the colorized frames into a single video file by using a video processing library to sequence the frames, combine them with the original audio data, and save the resulting video file.
[0930] Step 6:
[0931] The server saves the completed colorized video file in the server, generates a link so that the user can download it, and notifies the user of the generated link.
[0932] Step 7:
[0933] Users can download the colorized video file to their device using the download link provided by the server. When they click the link, the file is automatically saved in the download folder.
[0934] As a concrete example, if a user uploads a black-and-white video titled "Life in Ancient Rome," the video will be colorized through the steps described above. The user can then download the colorized version of "Life in Ancient Rome" and view it on their device. This process makes it possible to provide high-quality colorized video in a short amount of time.
[0935] Example 1
[0936] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0937] Conventional techniques for colorizing black-and-white footage require significant time and effort, requiring manual work by people with specialized skills and knowledge. Furthermore, the quality of the colorization is often inconsistent, resulting in unnatural coloring. This makes it difficult to unlock the value of historical and vintage films. Therefore, there is a need for an efficient system for automatically and realistically colorizing black-and-white footage.
[0938] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0939] In this invention, the server includes a means for uploading video files, a means for extracting and preprocessing frames from the video files, a means for colorizing the preprocessed frames using a generative AI model, a means for reconstructing the colorized frames and saving them as video files, and a means for providing users with a download link for the saved video files, thereby enabling realistic colorization of monochrome videos in a short amount of time.
[0940] "Video file" refers to video data stored in digital format.
[0941] "Server" refers to a computer system that provides services over a network.
[0942] "Uploading" refers to the act of transferring data from a local device to a server.
[0943] "Frame" refers to the individual still images that make up a video file.
[0944] "Preprocessing" refers to the preliminary preparation work carried out to improve data quality and for analysis.
[0945] A "generative AI model" refers to a program created using machine learning algorithms that learns patterns from data and generates new data.
[0946] "Colorization" refers to the process of adding color to black and white video or images.
[0947] "Reconstructing" refers to reorganizing the processed data into a new format.
[0948] "Storing" refers to recording and maintaining digital data in a storage device.
[0949] "Download link" refers to a URL for obtaining a digital file over a network.
[0950] "Unifying the resolution" refers to the process of unifying the number of pixels in an image or video according to a certain standard.
[0951] "Noise removal" refers to the process of removing unnecessary random fluctuations from image or video data.
[0952] "Contrast adjustment" refers to the process of adjusting the difference between light and dark to improve the visibility of images and videos.
[0953] "Pre-trained data" refers to an existing dataset that a machine learning model has used during its training process.
[0954] "Coloring" refers to adding appropriate color information to a black and white frame.
[0955] "Historical materials" refers to records and documents about past events and cultures.
[0956] "Realistic colors" refers to realistic and natural color representation.
[0957] This invention is a system for automatically colorizing monochrome movies and video works. This system operates by linking the user's device, a server, and a generative AI model. Specifically, it has the following steps and functions:
[0958] First, the user accesses the server's upload page through a web browser on their device. There, they select a monochrome video file stored on their local disk and upload it to the server. For example, they use the web browser's "file selection" dialog. The file selected by the user is sent to the server via an HTTP POST request.
[0959] When the server receives the uploaded video file, it uses the video processing library FFmpeg to split the video file into frames. Each frame undergoes preprocessing and then the following steps:
[0960] 1. Resolution unification: Resize the resolution of each frame to 720p.
[0961] 2. Noise Reduction: Use a Gaussian filter to remove noise in the video.
[0962] 3. Contrast Adjustment: Perform histogram equalization and adjust the contrast of the frame.
[0963] Once preprocessing is complete, the colorization process is performed using a generative AI model. The server inputs the preprocessed frames into the AI model, which then predicts and applies the optimal color for each frame. The generative AI model used is built using TensorFlow and PyTorch and is trained on a variety of data based on historical documents and literature. For example, based on the training data that "Roman buildings use yellowish stone," a yellowish color is applied to a building.
[0964] The colorized frames are reconstructed again using FFmpeg, and all frames are joined together to form a continuous video file. At this time, if there is audio data from the original video, this is also integrated at the same time. This completes the colorized video file. Specifically, the following FFmpeg command is used to join the frames:
[0965] ffmpeg -r 24 -i frame%04d_colored.png -i original_audio.mp3 -c:v libx264 -c:a aac output_colored_video.mp4
[0966] Finally, the server stores the reconstructed color video file and generates a downloadable link that can be provided to the user, who can use it to download the colorized video file to their device and view it.
[0967] As a concrete example, consider the case where a user uploads a monochrome video of "life in ancient Rome" to a server. The user accesses the server's upload page using a web browser on their device and uploads the "life in ancient Rome" video file. The server then divides the video into frames and preprocesses each frame. The preprocessed frames are then fed into a generative AI model for colorization. For example, each frame is colored with the optimal color based on historical literature on the colors of Roman architecture and clothing. The colorized frames are then reconstructed to generate the final color video file. The user can then download the colorized video file to their device using the provided download link and watch it. This system is expected to rapidly achieve realistic colorization of monochrome video, significantly improving the visual value of historical video works.
[0968] The above is a specific embodiment of the present invention, and the operation and configuration of each part of the system have been specifically described. This system enables realistic and efficient colorization of monochrome footage, significantly reducing the time required compared to traditional manual colorization. Furthermore, the automated process ensures consistently high colorization quality, significantly enhancing the visual appeal of historical footage and vintage films.
[0969] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0970] Step 1:
[0971] The user launches a web browser on their device and accesses the server's upload page. Specifically, they enter the server's upload page URL (e.g., https: / / example.com / upload) in the browser's URL bar. When the page appears, they click the "Choose File" button, select a monochrome video file (e.g., roman_life.mp4) from their local disk, and press the "Upload" button. The input data is a monochrome video file, and the output is a video file uploaded to the server.
[0972] Step 2:
[0973] The server receives the uploaded video file. Then, it splits the video file into frames using FFmpeg. Specifically, it executes the following command:
[0974] ffmpeg -i roman_life.mp4 frame%04d.png
[0975] As a result, the input data is a video file, and the output data is the divided frame images (e.g., frame0001.png, frame0002.png, ...).
[0976] Step 3:
[0977] The server performs pre-processing for each divided frame, which includes the following specific steps:
[0978] Resolution uniformity: To resize each frame to 720p, for example, run the following command:
[0979] ffmpeg -i frame%04d.png -vf scale=1280:720 frame%04d_resized.png
[0980] Denoising: To remove noise using a Gaussian filter, run the following command:
[0981] ffmpeg -i frame%04d_resized.png -vf "gblur=sigma=2" frame%04d_denoised.png
[0982] Contrast adjustment: To perform histogram equalization, run the following command:
[0983] ffmpeg -i frame%04d_denoised.png -vf "eq=contrast=1.5" frame%04d_processed.png
[0984] The input data is a divided frame image, and the output data is a frame image that has been preprocessed.
[0985] Step 4:
[0986] The server inputs the preprocessed frames into a generative AI model for colorization. Specifically, a generative AI model built with TensorFlow or PyTorch is used. The prompt input to the model is in the form of "Please apply appropriate colors based on the frame image I will provide." The input data is the preprocessed frame image, and the output data is the colorized frame image.
[0987] Step 5:
[0988] The server then reconstructs the colorized frames and combines them into a single video file by running the following FFmpeg command:
[0989] ffmpeg -r 24 -i frame%04d_colored.png -i original_audio.mp3 -c:v libx264 -c:a aac output_colored_video.mp4
[0990] The input data are colorized frame images and the original audio data, and the output data is a reconstructed color video file.
[0991] Step 6:
[0992] The server stores the reconstructed color video file and generates a download link, which is provided to the user, who uses it to download the colorized video file to their device. The input data is the reconstructed color video file, and the output data is the download link.
[0993] (Application example 1)
[0994] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0995] Today, there is a growing need to colorize monochrome films and photographs to breathe new life into them. However, traditional colorization methods are time-consuming, costly, and often require specialized knowledge. This creates a need for a system that allows anyone to easily convert monochrome images to color. There is also a need for this functionality to be easily accessible anywhere via mobile devices such as smartphones and tablets.
[0996] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0997] In this invention, the server includes means for uploading video files to the server, means for extracting and preprocessing frames from the video files, means for colorizing the preprocessed frames using a generative AI model, means for reconstructing the colorized frames and saving them as video files, means for providing users with download links for the saved video files, and means for uploading the video files via a smartphone application and making the results available for download. This allows users to easily colorize monochrome videos using their smartphones and quickly obtain the results.
[0998] A "video file" is a digital file containing visual and audio information.
[0999] A "server" is a dedicated computer system that provides services to other computers over a computer network.
[1000] A "frame" is a unit of still images that make up a video file.
[1001] "Preprocessing" refers to processing performed before data analysis or processing, and is an operation performed to improve the quality of the data.
[1002] A "generative AI model" is a trained computational model that uses artificial intelligence technology to analyze data and output results.
[1003] "Colorization" is the process of adding color to monochrome images or videos.
[1004] "Reconstruction" is the process of reassembling decomposed or preprocessed data back into its original form.
[1005] A "download link" is a URL that allows a user to obtain a file via the Internet.
[1006] "User" means a person or entity that uses the system or service.
[1007] A "smartphone application" is a software program that runs on a mobile device.
[1008] This invention is a system for automatically colorizing monochrome video works, and operates in cooperation with a user, a server, and a terminal. Specific embodiments of the invention are described below.
[1009] First, the user launches the application on their smartphone and selects the monochrome video file they want to colorize. This video file is then uploaded to the server via the smartphone application. The uploaded video file is then saved on the server, and the next process begins.
[1010] When the server receives the video file, it divides it into frames and preprocesses each frame. This preprocessing includes unifying the resolution, removing noise, and adjusting the contrast. This is done using software libraries such as OpenCV. Each preprocessed frame is then colorized using a generative AI model. This generative AI model has been trained in advance on a large amount of data, enabling highly accurate colorization.
[1011] The colorized frames are then reconstructed on the server and converted back into the original video file format. The reconstructed color video file is stored on the server, and a download link is generated that users can access. By providing this link to users, they can download the colorized video file via their smartphone.
[1012] Through this process, users can easily colorize monochrome images using their smartphones and receive the results quickly.
[1013] As a concrete example, consider a scenario in which a user uploads a monochrome video titled "Video of an Ancient City." The user first selects and uploads the video file through a smartphone application. The server receives the video file, splits it into frames using OpenCV, standardizes the resolution, removes noise, and adjusts the contrast. Then, a generative AI model colorizes the video, assigning appropriate colors to each frame. The server then reconstructs the colorized frames and generates a download link to provide to the user. Using this link, the user can download and watch the colorized video.
[1014] Another example of a prompt is, "Please input a monochrome frame and generate appropriate colors based on that frame. For example, the building should be the color of stone, the sky should be blue, and the ground should be brown," and the AI model will generate appropriate colors.
[1015] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1016] Step 1:
[1017] A user launches a smartphone application, selects a monochrome video file, and uploads it. This operation causes the smartphone application to send the selected video file to the server. The input is the user's video file, and the output is the file uploaded to the server. At this time, an HTTP POST request is used to send the file to the server.
[1018] Step 2:
[1019] The server receives the uploaded video file and splits it into frames. The input is the video file uploaded to the server, and the output is the split frames. The server uses the OpenCV library to perform the specific operation of splitting the video file into frames.
[1020] Step 3:
[1021] The server preprocesses each frame. The input is a set of split frames, and the output is a set of preprocessed frames. Preprocessing includes unifying the resolution, removing noise, and adjusting the contrast. Specifically, OpenCV is used to unify the resolution of each frame to 720p, apply a noise removal filter, and adjust the contrast using histogram equalization.
[1022] Step 4:
[1023] The server inputs the preprocessed frames into a generative AI model for colorization. The input is a set of preprocessed frames, and the output is a set of colorized frames. The generative AI model applies the appropriate color to each frame based on pre-trained data. Specifically, it uses TensorFlow or PyTorch models to colorize each frame.
[1024] Step 5:
[1025] The server reconstructs the colorized frames and returns them to the original video file format. The input is a set of colorized frames, and the output is a reconstructed color video file. The server uses libraries such as OpenCV to perform the specific operations of successively combining each frame and reconstructing it into the original video format.
[1026] Step 6:
[1027] The server saves the reconstructed color video file and generates a download link for users to access. The input is the reconstructed color video file, and the output is the download link. The server then performs the specific operation of generating a URL for the saved video file and providing it to the user.
[1028] Step 7:
[1029] A user accesses a download link through a smartphone application and downloads a colorized video file. The input is the download link, and the output is the colorized video file on the smartphone. The user clicks on the generated link and performs the specific action of downloading the file using an HTTP GET request.
[1030] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1031] The present invention relates to a system that automatically colorizes monochrome movies and video works by combining an emotion engine with a system that adjusts colors and filters based on the user's emotions. Specific embodiments of the present invention will be described below.
[1032] System Overview
[1033] 1. Uploading video files
[1034] The user uploads monochrome video files to the server through a specified interface from their own terminal. The user accesses the server's upload page using a web browser, selects a local video file using a file selection dialog, and performs the upload operation.
[1035] 2. Frame extraction and preprocessing
[1036] When the server receives the uploaded video file, it divides the video data into frames. Each frame undergoes preprocessing to unify the resolution, remove noise, and adjust the contrast, making it suitable for subsequent AI processing.
[1037] 3. Analysis by Emotion Engine
[1038] An emotion engine installed on the server recognizes the user's emotions. The emotion engine collects the user's facial expression data, voice data, or biometric information, and analyzes this data to identify the user's emotions. For example, if the user shows a happy expression while watching, that emotion data is analyzed.
[1039] 4. Adjusting the colorization process
[1040] Based on the analysis results of the emotion engine, the colorization parameters for the preprocessed frames are adjusted. For example, if the user is emotional, warm colors are emphasized. The preprocessed frames are input into the generative AI model, which then colorizes them based on the adjusted parameters.
[1041] 5. Reconstructing the frame
[1042] The server then reassembles the colorized frames into a single video file, resulting in a complete colorized video file. The reconstruction process also integrates audio data, if necessary.
[1043] 6. Providing download links
[1044] The server stores the reconstructed color video file in the server and generates a link for the user to download it. The created link is notified to the user, who uses it to download the colorized video file to their terminal.
[1045] Specific examples
[1046] For example, consider a scenario in which a user uploads a black and white video entitled "Life in Ancient Rome." The user accesses the server's upload page through a web browser and uploads this video file.
[1047] The server receives the "Life in Ancient Rome" video and splits it into frames, then standardizes the resolution of each frame to 720p, denoises it, and adjusts the contrast, preparing the pre-processed frames.
[1048] If the emotion engine analyzes the user's facial expressions and voice and recognizes that the user has positive emotions toward the video, it adjusts the colorization parameters based on this emotion data to emphasize warm colors.
[1049] The pre-processed frames are then fed into a generative AI model, which then colorizes them based on the adjusted parameters. For example, each frame is colored with the optimal color based on historical documents about the colors of Roman architecture and clothing.
[1050] The colorized frames are then reassembled and saved as a single video file. Finally, users can download the colorized version of "Life in Ancient Rome" to their devices using a download link provided by the server.
[1051] The present invention is expected to provide realistic colorized images based on the user's emotions in a short time, significantly improving visual value and user experience.
[1052] The processing flow will be explained below.
[1053] Step 1:
[1054] Users upload monochrome video files to the server from their own devices through a designated interface. This operation is performed using a web browser, accessing the upload page, selecting a local video file from the file selection dialog, and pressing the upload button.
[1055] Step 2:
[1056] The server receives the uploaded video file and temporarily stores it in a specified directory on the server, before moving on to the next processing step.
[1057] Step 3:
[1058] The server splits the video file into frames, analyzes the video frame by frame using a video processing library, extracts each frame as an image file, and stores it in temporary storage.
[1059] Step 4:
[1060] The server performs pre-processing on each extracted frame. First, it resizes it to the specified resolution (e.g., 720p) to ensure uniformity. Then it applies a noise reduction filter to improve the image quality. It also performs contrast adjustments to ensure visual consistency across the frames.
[1061] Step 5:
[1062] The emotion engine installed on the server recognizes the user's emotions. It collects facial expressions and voice data from the video the user is watching, and analyzes this data to identify the user's emotions. The emotion engine uses facial recognition and voice analysis technology to determine the user's emotional state (happiness, surprise, sadness, etc.).
[1063] Step 6:
[1064] The server automatically adjusts the colorization parameters based on the analysis results of the emotion engine. For example, if the user is recognized as happy, the colors are set to emphasize warm colors. The parameters provided by the emotion engine become input data for the generative AI model.
[1065] Step 7:
[1066] The server inputs the preprocessed frames into the generative AI model based on the adjusted parameters. The AI model selects the optimal color for each frame and colorizes it. The model then references the trained dataset and runs an algorithm to generate realistic colors.
[1067] Step 8:
[1068] The server reconstructs the colorized frames into a single video file, sequences the colorized frames using a video processing library, combines them with the original audio data, and saves the resulting video file.
[1069] Step 9:
[1070] The server stores the generated color image file in the server and generates a link for the user to download it, and notifies the user of the generated link.
[1071] Step 10:
[1072] Users can download the colorized video file to their device using the download link provided by the server. When they click the link, the file is automatically saved in the download folder.
[1073] Specific examples
[1074] For example, if a user uploads a monochrome video titled "Life in Ancient Rome," the video will be colorized through the steps described above. If the emotion engine analyzes the user's facial expressions and voice data and determines that the user has a positive emotion toward the video, the colorization will be adjusted to emphasize warm colors. This process allows the user to download and watch "Life in Ancient Rome," with enhanced visual value and emotional empathy. The present invention is expected to provide realistic colorized video tailored to the emotions of individual users in a short time, significantly improving visual value and user experience.
[1075] Example 2
[1076] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1077] Conventional colorization technologies simply colorize monochrome images and are unable to reflect the user's emotions or visual preferences. As a result, there is a gap between the colorized image and the user's emotions, resulting in a degradation of the viewing experience. Furthermore, due to inconsistent colorization accuracy and quality, parts of the image can appear unnatural.
[1078] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1079] In this invention, the server includes means for uploading a video file to an information processing device, means for extracting and preprocessing frames from the video file, means for collecting and analyzing user emotional information, means for adjusting colorization parameters for the preprocessed frames based on the analysis results, means for colorizing the preprocessed frames using a generative AI model, means for reconstructing the colorized frames and saving them as video files, and means for providing the user with a download link for the saved video file.
[1080] This allows for the provision of more realistic and moving colorized images based on the user's emotions, improving the quality of the viewing experience.
[1081] An "information processing device" is a device for collecting, analyzing, storing, and transmitting data, and includes a server or computer system.
[1082] "Video file" is a digital file containing video data, including monochrome video uploaded by users and final colorized video files.
[1083] A "frame" is a unit of still images that make up a video, and is the basic unit into which a video file is divided when it is processed.
[1084] "Preprocessing" refers to the processing performed on video frames, including resolution unification, noise removal, contrast adjustment, etc.
[1085] "Emotion information" is data that indicates the user's emotional state, and includes facial expression data, voice data, biometric information, and the like.
[1086] An "emotion engine" is a system that analyzes a user's emotional information and identifies their emotional state.
[1087] "Colorization parameters" are setting values for adjusting the color of the frame, and are determined based on the emotion analysis results.
[1088] A "generative AI model" is an artificial intelligence model that colorizes video frames based on pre-trained data.
[1089] "Reconstruction" is the process of connecting the colorized frames in order to recreate the original video file.
[1090] "Download link" refers to a URL or hyperlink that allows a user to download a stored video file via the Internet.
[1091] "Noise reduction" is the process of removing unwanted noise from a video frame.
[1092] "Contrast adjustment" is the process of improving the visibility of an image by adjusting the brightness of the frame.
[1093] "Upload" is the act of a user sending data from their own terminal to a server.
[1094] The present invention relates to a system that automatically colorizes monochrome movies and video works by combining an emotion engine with a system that adjusts colors and filters based on the user's emotions. Specific embodiments of the present invention will be described below.
[1095] (1. Uploading video files)
[1096] The user accesses the upload page of the information processing device using a web browser. The user clicks the "Select File" button on the page and selects a monochrome video file from the local file system. Next, the user presses the upload button to send the video file to the server. The server stores the received video file in a database.
[1097] (2. Frame Extraction and Preprocessing)
[1098] The server passes the received video file to the video analysis module, which divides the video file into frames and sends each frame to the image processing module, which standardizes the resolution to 720p and performs preprocessing such as noise reduction and contrast adjustment. This generates frames suitable for subsequent processing.
[1099] (3. Analysis by Emotion Engine)
[1100] When viewing the video, the user allows the device's camera and microphone to be used. The server collects the user's facial expression data, voice data, and biometric information (e.g., heart rate) in real time. This data is passed to the emotion engine, which analyzes the user's emotions. For example, if the user smiles, the emotion engine analyzes it as "joy."
[1101] (4. Adjusting colorization processing)
[1102] The server's colorization module receives the emotion analysis results from the emotion engine and adjusts the colorization parameters for the preprocessed frames. For example, if the user is emotional, the colorization parameters are set to emphasize warm colors. These adjusted parameters are input as prompts to the generative AI model to colorize each frame. An example prompt is: "The user is emotional. Please use warmer colors."
[1103] (5. Reconstructing the frame)
[1104] The server then reconstructs the colorized frames into a single video file. The reconstruction module then concatenates the frames together to form a continuous video file. If necessary, audio data is also integrated at this stage.
[1105] (6. Providing a download link)
[1106] The server saves the reconstructed color video file in its own storage. It then automatically generates a link that allows the user to download the video file. The user is notified of this link, and can click it to download the colorized video file to their device.
[1107] Specific examples
[1108] For example, consider a scenario in which a user wants to upload a black-and-white video entitled "Life in Ancient Rome." The user uses a web browser to access the upload page of the information processing device and uploads this video file.
[1109] The server receives the "Life in Ancient Rome" video and splits it into frames. Then it standardizes the resolution of each frame to 720p, removes noise, and adjusts the contrast. The pre-processed frames are prepared.
[1110] If the emotion engine analyzes the user's facial expressions and voice and determines that the user has positive emotions toward the video, it adjusts the colorization parameters based on this emotion data to emphasize warm colors. Example prompt: "The user has positive emotions. Use warm colors."
[1111] The pre-processed frames are then fed into a generative AI model, which then colorizes them based on the adjusted parameters. For example, each frame is colored with the optimal color based on historical documents about the colors of Roman architecture and clothing.
[1112] The colorized frames are then reassembled and saved as a single video file. Finally, users can download the colorized version of "Life in Ancient Rome" to their devices using a download link provided by the server.
[1113] The present invention is expected to provide realistic colorized images based on the user's emotions, significantly improving visual satisfaction and user experience.
[1114] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1115] Step 1:
[1116] The user launches a web browser and accesses the upload page of the information processing device. The user clicks the "Select File" button and selects a monochrome video file from the local file system. The user then presses the upload button to send the video file to the server. The input is the video file on the user's local system, and the output is the video file saved on the server.
[1117] Step 2:
[1118] The server passes the received video file to the video analysis module. The video analysis module divides the video file into frames. Each frame is then sent to the image processing module. The image processing module standardizes the resolution to 720p and performs preprocessing such as noise reduction and contrast adjustment. The input is the uploaded video file, and the output is a set of preprocessed frames.
[1119] Step 3:
[1120] When viewing a video, the user allows the device's camera and microphone to be used. The server collects the user's facial expression data, voice data, and biometric information (e.g., heart rate) in real time. This data is passed to the emotion engine. The emotion engine analyzes the user's emotions and generates the analysis results. The input is the user's real-time biometric data, and the output is the user's emotion analysis results.
[1121] Step 4:
[1122] The server's colorization module receives the emotion analysis results from the emotion engine and adjusts the colorization parameters for the preprocessed frames. The emotion analysis results are provided as input, and a prompt sentence is generated based on the results. This prompt sentence is input into a generative AI model to colorize each frame. For example, if the user is emotional, the colorization parameters are set to emphasize warm colors. The input is the emotion analysis results, and the output is the prompt sentence with the adjusted colorization parameters.
[1123] Step 5:
[1124] The server uses a generative AI model to colorize the preprocessed frames using adjusted parameters. Specifically, the generative AI model adds color to each frame based on the prompt. For example, each frame is assigned an appropriate color based on historical documents about the colors of Roman architecture and clothing. The input is the preprocessed frames and the prompt, and the output is a set of colorized frames.
[1125] Step 6:
[1126] The server then reconstructs the colorized frames into a single video file. The reconstruction module concatenates the frames together to form a continuous video file. If necessary, audio data is also integrated at this stage. The input is a set of colorized frames, and the output is a reconstructed color video file.
[1127] Step 7:
[1128] The server saves the reconstructed color video file in its own storage. It then automatically generates a link that allows the user to download the video file. This link is notified to the user, who clicks the link to download the colorized video file to their device. The input is the reconstructed color video file, and the output is the download link.
[1129] (Application example 2)
[1130] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1131] Conventional systems for colorizing black-and-white movies and video works simply add color, and do not provide a personalized viewing experience based on the user's emotions. As a result, there is a problem that the viewing experience is limited because the color adjustment does not correspond to the viewer's emotions. To solve this problem, a system that can colorize in accordance with the user's emotions is needed.
[1132] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for uploading a video file to the information processing device, means for extracting and preprocessing frames from the video file, means for colorizing the preprocessed frames using a generative AI model, means for analyzing a user's emotions, means for adjusting colorization parameters based on the analysis results, means for reconstructing the colorized frames and saving them as a video file, and means for providing a download link for the saved video file. This enables real-time colorization based on the user's emotions.
[1133] "Video file" means a data file that stores the content of visual media in digital format.
[1134] An "information processing device" is a device that processes and manages data, and includes computer systems such as servers.
[1135] A "frame" is a unit of each still image that makes up a video.
[1136] "Preprocessing" refers to the process of preparing data for subsequent processing.
[1137] A "generative AI model" is a model trained using artificial intelligence to perform a specific task, in this case colorization.
[1138] "Means for analyzing emotions" refers to a processing method for identifying emotions based on the user's facial expressions, voice, biometric information, etc.
[1139] A "parameter" is a setting that adjusts the behavior of a system or model.
[1140] "Reconstruction" means reassembling divided data into a single piece of data.
[1141] "Download link" refers to the access means by which a user can save the required data on their device via the Internet.
[1142] The present invention relates to a system for automatically colorizing monochrome movies and video works, and adjusting colors and filters based on the user's emotions. This system is configured and operates as follows.
[1143] First, a user uploads a monochrome video file to the information processing device through a specified interface on their own terminal. The user accesses the upload page using a web browser, selects a local video file using a file selection dialog, and performs the upload operation.
[1144] Next, when the server receives the uploaded video file, it divides the video data into frames. Each frame undergoes preprocessing to unify the resolution, remove noise, and adjust the contrast, making it suitable for subsequent AI processing.
[1145] The preprocessed frames are then analyzed by an emotion analysis engine installed on the server to determine the user's emotional state. The emotion analysis uses the user's facial expression data, voice data, or biometric information. For example, the emotion analysis engine uses cloud services such as EmotionAPI or Amazon Rekognition.
[1146] Based on the analysis, the server generates colorization parameters for the preprocessed frames. A generative AI model (such as OpenAI's GPT-4 or DALL-E) is used to colorize each frame based on these parameters. For example, if the user shows a happy expression, warm colors will be emphasized.
[1147] The colorized frames are then reconstructed into a single video file. During the reconstruction process, audio data is also integrated if necessary. Finally, the colorized video file is stored on a server, and a link is generated so that the user can download it. The created download link is notified to the user, who can use it to download the colorized video file to their device and view it.
[1148] Below are some example prompts that can be used as input to a generative AI model:
[1149] prompt:
[1150] Colorizing black and white "Classic Movies" while analyzing the user's emotional state:
[1151] If the user's emotion is "joy," use warm colors (e.g., warm orange tones)
[1152] If the user is emotional, use soft pastel colors.
[1153] Match the colors of each scene to their historical equivalents
[1154] Detailed frame-by-frame colorization
[1155] This makes it possible to provide an optimal video experience that suits the user's emotions.
[1156] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1157] Step 1:
[1158] A user uploads a monochrome video file to the server through a web browser. The user accesses a web page, selects a video file from a local file selection dialog, and uploads it. The input of this operation is a local monochrome video file, and the output is a video file saved on the server.
[1159] Step 2:
[1160] The server splits the uploaded video file into frames. The server analyzes the video file and extracts each frame. The input is the video file, and the output is a collection of individual frames.
[1161] Step 3:
[1162] The server pre-processes each frame, unifying the resolution of each frame, denoising, and adjusting the contrast. The input is a set of frames, and the output is a set of pre-processed frames.
[1163] Step 4:
[1164] The server analyzes the user's emotions from the preprocessed frames. Using an emotion analysis engine, the server collects and analyzes the user's facial expression data, voice data, and biometric information. The input is the user's facial expression data, voice data, and biometric information, and the output is the user's emotional state.
[1165] Step 5:
[1166] The server generates colorization parameters based on the analysis results. Based on the generated emotion data, the generative AI model generates a prompt sentence, and sets the colorization parameters for each frame based on that prompt sentence. The input is the user's emotional state, and the output is the colorization parameters.
[1167] Step 6:
[1168] The server inputs the preprocessed frames into a generative AI model to perform colorization. The generative AI model (e.g., GPT-4 or DALL-E) colorizes each frame based on the adjusted parameters. The input is a set of preprocessed frames and colorization parameters, and the output is a set of colorized frames.
[1169] Step 7:
[1170] The server reconstructs the colorized frames and saves them as the final color video file. It performs inter-frame reconstruction and integrates audio data as needed. The input is a set of colorized frames and the audio data from the original video, and the output is a reconstructed color video file.
[1171] Step 8:
[1172] The server generates a download link for the reconstructed color video file and provides it to the user. The download link is sent to the user via email or notification. The input is the reconstructed color video file, and the output is the download link provided to the user.
[1173] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1174] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1175] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1176] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1177] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1178] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1179] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1180] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1181] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1182] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1183] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1184] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1185] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1186] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1187] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1188] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1189] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1190] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1191] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1192] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1193] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1194] The following is further disclosed regarding the above embodiment.
[1195] (Claim 1)
[1196] A means for uploading video files to the server;
[1197] A means for extracting and preprocessing frames from a video file;
[1198] a means for colorizing the preprocessed frames using a generative AI model; and
[1199] A means for reconstructing the colorized frames and saving them as video files;
[1200] means for providing a user with a download link for the stored video file;
[1201] A system including:
[1202] (Claim 2)
[1203] 10. The system of claim 1, wherein the preprocessing means comprises means for unifying the resolution of each frame, removing noise, and adjusting contrast.
[1204] (Claim 3)
[1205] The system of claim 1, wherein the generative AI model includes means for applying color based on learned data for each frame.
[1206] "Example 1"
[1207] (Claim 1)
[1208] A means for uploading video files to the server;
[1209] A means for extracting and preprocessing frames from a video file;
[1210] a means for colorizing the preprocessed frames using a generative AI model; and
[1211] A means for reconstructing the colorized frames and saving them as video files;
[1212] means for providing a user with a download link for the stored video file;
[1213] A system including:
[1214] (Claim 2)
[1215] 10. The system of claim 1, wherein the preprocessing means comprises means for unifying the resolution of each frame, removing noise, and adjusting contrast.
[1216] (Claim 3)
[1217] The system of claim 1, wherein the generative AI model includes means for applying color based on learned data for each frame.
[1218] (Claim 4)
[1219] 10. The system of claim 1, wherein said means for reconstructing includes means for combining colorized frames into a continuous video format and integrating audio data as needed.
[1220] (Claim 5)
[1221] 10. The system of claim 1, wherein the generative AI model includes means for applying realistic colors by learning from historical and document-based data.
[1222] "Application Example 1"
[1223] (Claim 1)
[1224] A means for uploading video files to the server;
[1225] A means for extracting and preprocessing frames from a video file;
[1226] a means for colorizing the preprocessed frames using a generative AI model; and
[1227] A means for reconstructing the colorized frames and saving them as video files;
[1228] means for providing a user with a download link for the stored video file;
[1229] A means to upload video files and make the results downloadable via a smartphone application;
[1230] A system including:
[1231] (Claim 2)
[1232] 10. The system of claim 1, wherein the preprocessing means comprises means for unifying the resolution of each frame, removing noise, and adjusting contrast.
[1233] (Claim 3)
[1234] The system of claim 1, wherein the generative AI model includes means for applying color based on learned data for each frame.
[1235] "Example 2: Combining Emotion Engines"
[1236] (Claim 1)
[1237] means for uploading a video file to an information processing device;
[1238] A means for extracting and preprocessing frames from a video file;
[1239] means for collecting and analyzing user emotional information;
[1240] means for adjusting colorization parameters for the preprocessed frames based on the analysis results;
[1241] a means for colorizing the preprocessed frames using a generative AI model; and
[1242] A means for reconstructing the colorized frames and saving them as video files;
[1243] means for providing a user with a download link for the stored video file;
[1244] A system including:
[1245] (Claim 2)
[1246] 10. The system of claim 1, wherein the preprocessing means comprises means for unifying the resolution of each frame, removing noise, and adjusting contrast.
[1247] (Claim 3)
[1248] The system of claim 1, wherein the generative AI model includes means for applying color based on learned data for each frame.
[1249] "Application example 2 when combining emotion engines"
[1250] (Claim 1)
[1251] means for uploading a video file to an information processing device;
[1252] A means for extracting and preprocessing frames from a video file;
[1253] a means for colorizing the preprocessed frames using a generative AI model; and
[1254] means for analyzing user emotions;
[1255] a means for adjusting colorization parameters based on the analysis results;
[1256] A means for reconstructing the colorized frames and saving them as video files;
[1257] a means for providing a download link for the stored video file;
[1258] A system including:
[1259] (Claim 2)
[1260] 10. The system of claim 1, wherein the preprocessing means comprises means for unifying the resolution of each frame, removing noise, and adjusting contrast.
[1261] (Claim 3)
[1262] 2. The system of claim 1, wherein the generative AI model includes a prompt based on the analysis result and means for coloring each frame based on learned data. [Explanation of symbols]
[1263] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for uploading video files to the server; A means for extracting and preprocessing frames from a video file; a means for colorizing the preprocessed frames using a generative AI model; and A means for reconstructing the colorized frames and saving them as video files; means for providing a user with a download link for the stored video file; A system including:
2. 2. The system of claim 1, wherein the preprocessing means includes means for unifying the resolution of each frame, removing noise, and adjusting contrast.
3. The system of claim 1 , wherein the generative AI model includes means for applying color based on learned data for each frame.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A