System
The system addresses inefficiencies in video data analysis by optimizing content recommendation, ad delivery, and monitoring through automated processes, enhancing user experience and reducing costs.
Patent Information
- Application Number
- JP2024115219
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-01-29
AI Technical Summary
Current systems face challenges in efficiently analyzing massive amounts of video data for content recommendation, ad delivery, and manual monitoring, which is time-consuming and costly, while lacking effective age filters.
A system that includes video ingestion, quality verification, metadata generation, content analysis, categorization, recommendation based on user characteristics, ad selection, automatic/manual monitoring determination, and age filtering to optimize video services.
This system provides high-quality, efficient video services by tailoring content and ads to users, reducing monitoring costs, and ensuring a safe viewing environment.
Smart Images

Figure 2026014222000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] The present invention aims to provide a system that efficiently summarizes videos posted by users and, based on the summaries, improves recommendation functions, optimizes ad delivery, reduces monitoring costs, and provides a safe usage environment. Current systems have problems: analyzing massive amounts of video data takes time and effort, making it difficult to recommend appropriate content, deliver ads, and apply age filters. Furthermore, manual video monitoring is required, which is costly. There is a need to solve these issues and provide high-quality, efficient video services. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems by the following means. The system includes a means for uploading videos from a terminal, a means for a server to check the quality of the received videos and generate metadata, a means for the server to analyze the video content and generate a summary, a means for the server to categorize the videos based on the generated summary and assign tags and keywords, a means for the server to add the videos to a recommendation list based on user characteristics, a means for the server to select and assign optimal advertisements based on user characteristics, a means for the server to determine whether automatic or manual monitoring is required based on the video content, and a means for the server to apply an age filter based on the user's estimated age. This dramatically improves the quality and efficiency of video services. This makes it possible to provide content and advertisements tailored to users, reduce monitoring costs, and provide a safe and convenient environment.
[0006] A "terminal" is a device operated by a user, and is used to generate, upload, view, and otherwise use content.
[0007] A "server" is a centralized computer system that communicates with terminals via a network and receives, analyzes, stores, and distributes data.
[0008] "Video" is digital content that contains continuous video and audio, and is a media format that provides information through sight and hearing.
[0009] "Upload" refers to the operation of sending data from a terminal to a server and storing the data on the server.
[0010] "Quality verification" refers to the process of inspecting the format and playback status of received video data to determine whether it meets the quality standards.
[0011] "Metadata" refers to data that indicates information about a video (e.g., title, description, file size, playback time, etc.), rather than the content of the video itself.
[0012] "Analysis" is the process of examining the content of a video in detail and extracting important elements and patterns.
[0013] A "summary" is a sentence or data that concisely summarizes the main content or essence of a video.
[0014] "Categorization" refers to the process of classifying videos into specific categories or groups.
[0015] "Tags" are keywords or identifiers used to describe the content of a video.
[0016] "Keywords" are the main words or phrases that describe the content of your video.
[0017] "Recommendation" is a function that recommends the most suitable content to the user.
[0018] "User characteristics" refer to characteristics and attributes associated with individual users, such as their age, gender, interests, and viewing history.
[0019] "Advertising" means a commercial message displayed within a video with the purpose of raising awareness of a particular product or service.
[0020] "Automatic monitoring" refers to the process in which the system independently checks the video content and determines whether there are any problems.
[0021] "Manual monitoring" refers to the process of humans reviewing the content of videos and determining whether they are appropriate.
[0022] An "age filter" is a mechanism that restricts viewing of content that is not suitable for a certain age group. [Brief explanation of the drawings]
[0023] [Figure 1]1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0024] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0025] First, the terms used in the following description will be explained.
[0026] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0027] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0028] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0029] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0030] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0031] [First embodiment]
[0032] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0033] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0034] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0035] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0036] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0037] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0038] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0039] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0040] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0041] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0042] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0043] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0044] This invention is a system that efficiently summarizes videos posted by users and, based on the summaries, enables recommendation functions, advertisement distribution, reduced monitoring costs, and a safe usage environment. This system is built on communications between terminals, servers, and users. Below, the processing of the system's program is explained in natural language, with specific examples.
[0045] Video Ingestion and Analysis
[0046] 1. The device sends the video uploaded by the user to the server.
[0047] Device: The user uploads the video file and its meta information (title, description, etc.) to the server from a device such as a smartphone or PC.
[0048] 2. The server receives the video and performs the ingest process.
[0049] Server: Checks the format of the video file and converts it to a standard format if necessary. It also checks the video quality and generates metadata (file size, duration, etc.).
[0050] 3. The server uses AI technology to analyze the video content and generate a summary.
[0051] Server: Analyzes video frames to identify key scenes and important objects. Also, analyzes audio to identify speakers and extract important phrases. Then, it generates a video summary based on this information.
[0052] Categorizing the video
[0053] 1. The server categorizes the video based on the generated summary.
[0054] Server: Classifies videos into specific categories based on tags and keywords extracted from the summaries, and assigns these tags and keywords to the video data.
[0055] 2. The server creates a recommendation list suitable for each user.
[0056] Server: Based on the user's viewing history and interests, the server selects the most suitable videos and generates a recommendation list. This list is sent to the user's device.
[0057] Video Ad Optimization
[0058] 1. The server selects video ads based on user characteristics.
[0059] Server: Selects the most suitable advertisement based on the user's characteristics such as age, gender, and interests. Also sets the advertisement meta information and playback timing.
[0060] 2. The server delivers ads within the video.
[0061] Server: Evaluates the relevance of the video content and inserts ads at the appropriate time. When a user watches a video, relevant ads are displayed.
[0062] Determining what to monitor
[0063] 1. The server determines whether to monitor the video based on the summary.
[0064] Server: Detects whether the video summary contains inappropriate content (e.g. violence, adult content, etc.), performs a risk assessment, and determines whether manual supervision is required.
[0065] 2. The server will tag the data for manual monitoring as needed.
[0066] Server: Videos that are deemed safe by automatic monitoring are exempt from manual monitoring. If manual monitoring is deemed necessary, the monitoring team is notified.
[0067] Applying an age filter
[0068] 1. The server applies an age filter based on the user's estimated age.
[0069] Server: Estimates the user's age from their registration information and activity history. Detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[0070] 2. The server filters inappropriate content.
[0071] Server: Content that is not appropriate for the specified age group will not be played on the user's device, thereby preventing minors from accessing inappropriate content.
[0072] Specific examples
[0073] Example 1: User-Submitted Pet Videos
[0074] 1. The user (device) uploads a video of their dog playing at home.
[0075] 2. The server receives the video, performs quality checks, and generates metadata.
[0076] 3. The server analyzes the video and generates keywords such as "pet," "dog," and "play."
[0077] 4. The server uses these keywords to recommend pet videos relevant to the user.
[0078] Example 2: Fashion-related video ads
[0079] 1. The server identifies that the user has fashion-related interests.
[0080] 2. The server selects fashion-related video ads and inserts them into the video at the appropriate time.
[0081] 3. The device displays relevant ads while you watch fashion videos.
[0082] Example 3: Reducing monitoring costs
[0083] 1. The server verifies the summary of the posted video for inappropriate content.
[0084] 2. If the server determines that there is no problem, the video will be automatically made public and will no longer be subject to manual monitoring.
[0085] This system enables the provision of content tailored to user characteristics, optimization of advertising revenue, and reduction of human monitoring costs.
[0086] The processing flow will be explained below.
[0087] Step 1:
[0088] A user uses a terminal to upload a video file and its meta information (title, description, etc.) to a server.
[0089] Step 2:
[0090] The server checks the format of the video file received and converts it to a standard format if necessary, checks the video quality, and generates metadata (file size, playback time, etc.).
[0091] Step 3:
[0092] The server begins analyzing the video content using AI technology, specifically by analyzing the frames of the video and identifying key scenes and important objects.
[0093] Step 4:
[0094] The server performs audio analysis to identify speakers and extract key phrases, and generates a video summary based on the results of this analysis.
[0095] Step 5:
[0096] Based on the generated summary, the server assigns tags and keywords to the video and classifies it into a specific category, thereby categorizing the video.
[0097] Step 6:
[0098] The server selects appropriate videos based on the user's viewing history and interests, and generates a recommendation list for the user. This list is sent to the device.
[0099] Step 7:
[0100] The server selects the most suitable video ad based on user characteristics (age, gender, interests, etc.) and also sets the ad's meta information and playback timing.
[0101] Step 8:
[0102] The server evaluates the relevance of the video content and inserts ads at the appropriate time, and displays relevant ads when the device is watching the video.
[0103] Step 9:
[0104] Based on the summary, the server determines whether the video is suitable for automatic monitoring or requires manual monitoring. Inappropriate content (e.g., violence, adult content, etc.) is detected.
[0105] Step 10:
[0106] The server performs a risk assessment and, if it determines that there are no problems with automated monitoring, it makes the video public. If it determines that manual monitoring is necessary, it notifies the monitoring team.
[0107] Step 11:
[0108] The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the summary.
[0109] Step 12:
[0110] The server filters content that is inappropriate for the specified age and prevents the video from being played on the device, thereby preventing minors from accessing inappropriate content.
[0111] Example 1
[0112] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0113] Modern video streaming services require efficient analysis and classification of the vast amount of videos uploaded by users, and the provision of appropriate recommendations and advertisements. However, manual video analysis and recommendation generation requires a great deal of effort and time, resulting in high operational costs, making efficient and automated processing methods necessary. Automation is also required for providing personalized advertisements based on user characteristics, filtering content suitable for underage users, and monitoring video content. An efficient and reliable system is needed to solve these challenges.
[0114] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0115] In this invention, the server includes a means for uploading videos from a terminal, a means for converting the format of the received video, checking the quality, and generating metadata, a means for analyzing the content of the video using a generation AI model and generating a summary, a means for categorizing the video based on the generated summary and assigning tags and keywords, a means for adding the video to a recommendation list based on user characteristics, a means for selecting and delivering optimal advertisements based on user characteristics, a means for determining whether automatic or manual monitoring is required based on the video content, and a means for applying an age filter based on the estimated age of the user. This enables efficient analysis of videos uploaded by users, provision of appropriate recommendations and advertisements, automated content monitoring, and age-appropriate filtering.
[0116] A "terminal" is a device used by a user, such as a smartphone or computer, used to upload videos.
[0117] The "server" is a central device that receives, analyzes, and processes video data over the network, and plays a central role in this system.
[0118] A "video" is a file containing a series of video and audio files uploaded by a user.
[0119] "Format conversion" is the process of converting received video files into a standard format.
[0120] "Quality check" is the process of checking the quality of the received video, such as its resolution and playback time.
[0121] "Metadata" refers to various information related to a video file (e.g., file size, playback time, resolution, etc.).
[0122] A "generative AI model" is an artificial intelligence technology used to analyze video content and generate summaries.
[0123] "Video analysis" is the process of analyzing video frames and audio to identify important scenes and objects.
[0124] "Summary" refers to a short summary of the main content extracted from the results of video analysis.
[0125] "Tags" refer to keywords or labels that are added to indicate the content of a video.
[0126] "Keywords" refer to words or short phrases related to the video content.
[0127] A "recommendation list" is a list of videos selected based on the user's characteristics and interests.
[0128] "Advertisement" refers to information including product promotions and service information, which is inserted into videos so that users can view them.
[0129] "Automatic monitoring" is the process of automatically checking video content to detect inappropriate content.
[0130] "Manual monitoring" is a manual monitoring process carried out by specialized staff when automatic monitoring is difficult to judge.
[0131] "Estimated age" is the age calculated based on the user's registration information and activity history.
[0132] An "age filter" is a mechanism that provides only appropriate content based on the user's estimated age.
[0133] This invention is a system that efficiently summarizes videos posted by users and provides recommendation functions, advertisement distribution, reduced monitoring costs, and a safe usage environment based on the summaries. This system is built on communications between terminals, a server, and users.
[0134] Video Ingestion and Analysis
[0135] When a user uploads a video to a server from a device, the user uses a device such as a smartphone or PC. Specifically, the user sends the video file and its meta information (title, description, etc.) to the server through an interface provided by the device.
[0136] The server receives video files uploaded by users. The received video files are first converted into a different format. Software such as ffmpeg is used for this process. Next, the quality of the video is checked and metadata (file size, resolution, playback time, etc.) is generated based on that quality.
[0137] The server uses a generative AI model to analyze the video content and generate a summary. Specifically, it uses an object detection model such as YOLO (You Only Look Once) to identify key scenes and important objects, and then performs audio analysis using the Google Cloud Speech-to-Text API. Then, based on this information, it uses a natural language generation model (e.g., GPT-3) to generate a video summary.
[0138] Video categorization and recommendations
[0139] The server categorizes the videos based on the generated summaries. The videos are classified into specific categories based on tags and keywords extracted from the summaries. The classified videos are then assigned tags and keywords to improve searchability and filtering capabilities.
[0140] The server then generates a recommendation list based on the user's characteristics. Taking into account the user's viewing history and interests, the server selects the most suitable videos to create the recommendation list. This recommendation list is then sent to the user's device, where the user can watch the recommended videos through an interface.
[0141] Ad optimization and delivery
[0142] The server selects the most suitable advertisement based on the video summary and user characteristics. Specifically, it selects advertisements based on the user's characteristics such as age, gender, and interests, and sets the meta information and playback timing. The server references a database provided by the advertiser to select advertisements.
[0143] The server delivers advertisements at appropriate times within the video. For example, while watching a video about pets, it displays an advertisement for related pet food. This can be expected to have a high advertising effect.
[0144] Determining who should be monitored and applying age filters
[0145] The server determines whether a video should be subject to automated or manual monitoring based on the summary. The video summary is used to detect whether it contains inappropriate content and perform a risk assessment. Based on this, the process of determining whether to subject the video to automated or manual monitoring uses natural language processing technology.
[0146] Furthermore, the server applies an age filter based on the user's estimated age. The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the video summary. Inappropriate content is filtered so that it cannot be played on devices belonging to underage users. Based on these filtering rules, only appropriate videos are provided.
[0147] Specific examples
[0148] Example 1: User-Submitted Pet Videos
[0149] A user uploads a video of their dog playing at home. The server receives the video, checks the quality, and generates metadata. The server then analyzes the video and generates keywords such as "pet," "dog," and "play." The server then recommends relevant pet videos to the user based on these keywords.
[0150] Example 2: Fashion-related video ads
[0151] The server identifies the user's interest in fashion. The server then selects fashion-related video ads and inserts them into the video at the appropriate time. The relevant ads are then displayed on the user's device while they are watching the fashion video.
[0152] Example 3: Reducing monitoring costs
[0153] The server verifies the summary of the posted video to see if it contains inappropriate content. If it determines there are no problems, it automatically publishes the video and removes it from the scope of human monitoring.
[0154] Prompt Sentence Examples
[0155] "Recommend other videos related to this pet video."
[0156] "Insert fashion-related ads at the best possible time."
[0157] Please determine if this video is inappropriate.
[0158] This system makes it possible to provide content tailored to user characteristics, optimize advertising revenue, and reduce manual monitoring costs.
[0159] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0160] Step 1:
[0161] The device sends the video uploaded by the user to the server. The user clicks the upload button through the device interface, selects the video file (e.g., mp4, avi) and its meta information (title, description), and sends it to the server.
[0162] Input: Video file and meta information
[0163] Output: Video file and meta information sent to the server
[0164] Step 2:
[0165] The server receives the video, converts the format, checks the quality, and generates metadata. The server checks the format of the received video file and converts it to a standard format (e.g., mp4). This process uses tools such as ffmpeg. It also checks the quality of the video, such as resolution, playback time, and file size, and generates this information as metadata.
[0166] Input: Video file and meta information sent to the server
[0167] Output: Format converted video file and generated metadata
[0168] Step 3:
[0169] The server uses a generative AI model to analyze the video content and generate a summary. Specifically, it uses an object detection model such as YOLO to analyze video frames and identify key scenes and important objects. It also performs audio analysis using the Google Cloud Speech-to-Text API to identify speakers and extract key phrases. It then uses this data to generate a video summary using a natural language generation model (e.g., GPT-3).
[0170] Input: Format converted video file and metadata
[0171] Output: Video summary
[0172] Step 4:
[0173] The server categorizes the video based on the generated summary and assigns tags and keywords. Based on the information obtained from the summary, the server classifies the video into an appropriate category (e.g., pets, sports, entertainment), and assigns tags and keywords related to that category to the video data.
[0174] Input: Video summary
[0175] Output: Categorized videos with tags and keywords
[0176] Step 5:
[0177] The server adds videos to a recommendation list based on user characteristics. The server analyzes the user's viewing history and interests, selects the most suitable videos, and generates a recommendation list. This list is sent to the user's device.
[0178] Input: User's viewing history and interests, as well as categorized videos, tags, and keywords
[0179] Output: Recommendation list based on user characteristics
[0180] Step 6:
[0181] The server selects and delivers the most suitable advertisement based on the user's characteristics. The server selects the appropriate advertisement based on the user's characteristics information such as age, gender, and interests. Specifically, it retrieves the advertisement that best suits the user's characteristics from the advertisement database and sets its meta information and playback timing.
[0182] Input: User characteristics (age, gender, interests, etc.) and advertising database
[0183] Output: Optimal advertisement based on user characteristics
[0184] Step 7:
[0185] The server determines whether the video should be monitored automatically or manually based on its content. The server then checks the video summary for inappropriate content and assesses the content risk. This process uses natural language processing technology.
[0186] Input: Video summary
[0187] Output: Judgment results of automatic or manual monitoring
[0188] Step 8:
[0189] The server applies an age filter based on the user's estimated age. The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords from the video summary. Inappropriate content is filtered so that it cannot be played on devices belonging to minors.
[0190] Input: User registration information, activity history, video summary
[0191] Output: Age-filtered video list
[0192] (Application example 1)
[0193] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0194] In recent years, with the proliferation of video content, a huge amount of videos has been uploaded to online platforms. However, there is a lack of systems that can efficiently summarize videos and recommend appropriate content to users. Furthermore, manually monitoring video content and filtering inappropriate content is extremely time-consuming and costly. Furthermore, optimization of advertisements based on user characteristics is insufficient, and efficient advertisement delivery is required. To address these issues, a system is needed that automates video analysis and recommendation, advertisement optimization, and content monitoring, while providing a safe viewing environment.
[0195] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0196] In this invention, the server includes a means for uploading videos from a terminal, a means for checking the quality of the received videos and generating metadata, a means for analyzing the video content and generating a summary, a means for categorizing the videos based on the generated summary and assigning tags and keywords, a means for adding the videos to a recommendation list based on user characteristics, a means for selecting and assigning optimal advertisements based on user characteristics, a means for determining whether automatic or manual monitoring is required based on the video content, a means for applying an age filter based on the estimated age of the user, a means for saving extracted frame images and transcribing and analyzing audio, a means for providing a function for recommending similar content based on responses from the server, and a means for inserting advertisements and optimizing based on user characteristics as part of a content distribution service. This enables efficient video summarization and recommendation functions, optimal advertisement distribution, reduced monitoring costs, and a safe usage environment.
[0197] - "Device" refers to the device to which the video is uploaded, such as a smartphone or computer.
[0198] "Server" refers to the computing system that processes the received video, checks the quality, generates metadata, and analyzes the video content.
[0199] "Video content analysis" refers to the process of analyzing video frames and audio data to extract summaries, important scenes, and objects.
[0200] "Generating a summary" refers to extracting the main points from the analyzed video data and summarizing them concisely.
[0201] "Categorizing videos" refers to classifying videos into specific categories based on the generated summary and assigning relevant tags and keywords.
[0202] "Add to recommendation list" refers to selecting suitable videos based on the user's viewing history and interests and adding them to the display list.
[0203] "Selecting and delivering ads" refers to selecting the most suitable ads based on user characteristics and inserting them while the video is playing.
[0204] "Determine whether to use automated or manual monitoring" refers to determining whether a video contains inappropriate content based on a summary of the video and deciding whether to use automated or, if necessary, manual monitoring.
[0205] "Applying age filters" refers to restricting the content available to a user based on their estimated age.
[0206] "Save frame images" means saving still images extracted from a video, usually at 1-second intervals.
[0207] "Audio transcription and analysis" refers to the process of converting video audio into text data and extracting key phrases and keywords.
[0208] Providing a "similar content recommendation function" refers to a function that automatically suggests relevant videos based on a user's viewing history and characteristics.
[0209] "Ad insertion and optimization based on user characteristics" refers to inserting ads into videos at the appropriate time and providing the optimal advertising plan based on the user's characteristics.
[0210] This invention provides a system that efficiently summarizes videos posted by users and, based on the summaries, enables recommendation functions, advertisement distribution, reduced monitoring costs, and a safe usage environment. Detailed modes for carrying out the invention are described below.
[0211] Uploading videos from your device
[0212] Users upload video files to the server using devices such as smartphones or PCs. At this time, the video file and its meta information (title, description, etc.) are also sent.
[0213] Server receives and processes video
[0214] The server checks the quality of the received video files, converts them to a standard format if necessary, and generates metadata (file size, playback time, etc.) for the video.
[0215] Video content analysis and summary generation
[0216] The server analyzes the video frames to identify key scenes and important objects, and performs audio analysis to identify speakers and extract key phrases. Based on this information, a video summary is generated. This process includes saving the captured frame images and transcribing and analyzing the audio extracted from the video.
[0217] Video categorization and recommendations
[0218] The server categorizes the videos based on the generated summaries and assigns tags and keywords. Based on the user's characteristics, it selects the most suitable videos and generates a recommendation list. Based on the response from the server, a recommendation function for similar content is provided.
[0219] Ad optimization and insertion
[0220] The server selects the most suitable advertisement based on the user's characteristics and inserts it at a timing relevant to the video content. As part of the content distribution service, it performs advertisement insertion and optimization based on user characteristics.
[0221] Video monitoring and age filtering
[0222] The server determines whether to monitor the content based on the video summary, automatically or manually, and applies an age filter based on the user's estimated age to filter out inappropriate content.
[0223] Hardware and software used
[0224] The following hardware and software are used to implement this system.
[0225] Hardware: Smartphones, PCs
[0226] Software: MoviePy (video editing library), Vosk (voice recognition library), OpenCV (image processing library)
[0227] Specific examples
[0228] Example 1: User-Submitted Pet Videos
[0229] When a user uploads a video of their dog playing at home, the server receives the video, checks the quality, and generates metadata. It then analyzes the video to generate keywords such as "pet," "dog," and "play," and recommends related pet videos.
[0230] Example 2: Fashion-related video ads
[0231] The server identifies that the user is interested in fashion, selects fashion-related video advertisements, and inserts them into the videos at the appropriate time. The relevant advertisements are displayed while the user is watching the fashion video.
[0232] Prompt Sentence Examples
[0233] text
[0234] Analyzes the user-uploaded video "example_video.mp4" and generates a summary. Then, recommends similar videos and inserts relevant ads. Finally, applies supervision and age filters.
[0235] In this way, the present invention makes it possible to provide efficient video summarization and recommendation functions, optimal advertisement distribution, reduced monitoring costs, and a safe usage environment.
[0236] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0237] Step 1:
[0238] A user uses a terminal to upload a video file to a server.
[0239] Input: Video file, meta information (title, description, etc.)
[0240] Output: Video files and meta information are saved on the server
[0241] Specific operation: A user selects a video file and presses the upload button on their smartphone or PC. The device then sends the selected video file and its meta information to the server as an HTTP request.
[0242] Step 2:
[0243] The server checks the quality of the received video file and generates metadata.
[0244] Input: Uploaded video file, meta information
[0245] Output: Video quality evaluation results, metadata (file size, playing time, etc.)
[0246] What it does: The server checks the format of the video file and converts it to a standard format if necessary. It also evaluates the video quality and generates metadata such as file size, duration, and resolution.
[0247] Step 3:
[0248] The server performs frame analysis and audio analysis of the video content and generates a summary.
[0249] Input: Video file
[0250] Output: Video summary, list of important scenes and objects, audio transcription
[0251] Specific operations: Extract key scenes from each frame of a video using MoviePy and OpenCV. Transcribe audio data using Vosk and extract key phrases and keywords. Combine these to generate a video summary.
[0252] Step 4:
[0253] The server categorizes the video based on the generated summary and assigns tags and keywords.
[0254] Input: Video summary
[0255] Output: Video data with tags and keywords
[0256] How it works: The server classifies videos into specific categories based on phrases and keywords extracted from the summaries, automatically assigning tags such as "pets," "dogs," and "play."
[0257] Step 5:
[0258] The server generates a recommendation list based on user characteristics.
[0259] Input: User viewing history, interests, summary data
[0260] Output: Recommendation list for each user
[0261] How it works: The server analyzes the user's past viewing history and interests (e.g., pet videos) and selects relevant videos from the summary data, thereby generating a customized recommendation list for each user.
[0262] Step 6:
[0263] The server selects the most suitable advertisement based on user characteristics and inserts it at the appropriate time within the video.
[0264] Input: User characteristics, video summary, relevant advertising data
[0265] Output: Video with ads
[0266] Specific operation: The server selects the most suitable advertisement based on the user's characteristics such as age, gender, and interests, and inserts the selected advertisement at the relevant timing in the video (e.g., scene change).
[0267] Step 7:
[0268] The server determines which videos to monitor based on the video summary.
[0269] Input: Video summary
[0270] Output: Judgment result of whether the target is automatic or manual monitoring
[0271] How it works: The server evaluates the summary and determines whether it contains inappropriate content. If the automated monitoring finds no problems, the video is left public. If manual monitoring is deemed necessary, the monitoring team is notified.
[0272] Step 8:
[0273] The server applies an age filter based on the user's estimated age.
[0274] Input: User registration information, activity history, video summary
[0275] Output: Filtered video content
[0276] Specific operation: The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords from the video summary. It then filters the video to prevent underage users from accessing inappropriate content.
[0277] Examples of prompt statements
[0278] Analyzes the user-uploaded video "example_video.mp4" and generates a summary. Then, recommends similar videos and inserts relevant ads. Finally, applies supervision and age filters.
[0279] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0280] This invention is a system that efficiently summarizes videos posted by users, improving recommendation functionality, optimizing ad delivery, reducing monitoring costs, and providing a safe usage environment. Furthermore, by combining it with an emotion engine that analyzes user emotions, more advanced recommendations and ad delivery become possible. This system is built on communications between terminals, servers, and users. Below, we explain the system's program processing in natural language, including specific examples.
[0281] Video Ingestion and Analysis
[0282] 1. The device sends the video uploaded by the user to the server.
[0283] Device: The user uploads the video file and its meta information (title, description, etc.) to the server from a device such as a smartphone or PC.
[0284] 2. The server receives the video and performs the ingest process.
[0285] Server: Checks the format of the video file and converts it to a standard format if necessary. It also checks the video quality and generates metadata (file size, duration, etc.).
[0286] 3. The server uses AI technology to analyze the video content and generate a summary.
[0287] Server: Analyzes video frames to identify key scenes and important objects. Also, analyzes audio to identify speakers and extract important phrases. Then, it generates a video summary based on this information.
[0288] Categorizing the video
[0289] 1. The server categorizes the video based on the generated summary.
[0290] Server: Classifies videos into specific categories based on tags and keywords extracted from the summaries, and assigns these tags and keywords to the video data.
[0291] 2. The server creates a recommendation list suitable for each user.
[0292] Server: Based on the user's viewing history and interests, the server selects the most suitable videos and generates a recommendation list. This list is sent to the user's device.
[0293] Emotion analysis and recommendations using an emotion engine
[0294] 1. The server uses an emotion engine to analyze the user's emotions.
[0295] Server: Analyzes emotions based on the user's facial expressions, tone of voice, viewing history, etc.
[0296] 2. The server reflects the results of the sentiment analysis in its video recommendations.
[0297] Server: Updates the recommendation list taking into account emotional data to provide more of the content the user enjoys.
[0298] Video Ad Optimization
[0299] 1. The server selects video ads based on user characteristics.
[0300] Server: Selects the most suitable advertisement based on the user's age, gender, interests, and emotional analysis results. Also sets the advertisement meta information and playback timing.
[0301] 2. The server delivers ads within the video.
[0302] Server: Evaluates the relevance of the video content and inserts ads at the appropriate time. When a user watches a video, relevant ads are displayed.
[0303] Determining what to monitor
[0304] 1. The server determines whether to monitor the video based on the summary.
[0305] Server: Detects whether the video summary contains inappropriate content (e.g. violence, adult content, etc.), performs a risk assessment, and determines whether manual supervision is required.
[0306] 2. The server will tag the data for manual monitoring as needed.
[0307] Server: Videos that are deemed safe by automatic monitoring are exempt from manual monitoring. If manual monitoring is deemed necessary, the monitoring team is notified.
[0308] Applying an age filter
[0309] 1. The server applies an age filter based on the user's estimated age.
[0310] Server: Estimates the user's age from their registration information and activity history. Detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[0311] 2. The server filters inappropriate content.
[0312] Server: Content that is not appropriate for the specified age group is not played on the user's device, thereby preventing minors from accessing inappropriate content.
[0313] Specific examples
[0314] Example 1: User-Submitted Pet Videos
[0315] 1. The user (device) uploads a video of their dog playing at home.
[0316] 2. The server receives the video, performs quality checks, and generates metadata.
[0317] 3. The server analyzes the video and generates keywords such as "pet," "dog," and "play."
[0318] 4. The server uses these keywords to recommend pet videos relevant to the user.
[0319] 5. The server analyzes the user's emotions, detects whether the user is happy, and recommends more pet videos based on that.
[0320] Example 2: Fashion-related video ads
[0321] 1. The server identifies that the user has fashion-related interests.
[0322] 2. The server uses the emotion engine to confirm that the user has positive emotions toward the fashion content.
[0323] 3. The server selects fashion-related video ads and inserts them into the video at the appropriate time.
[0324] 4. The device displays relevant ads while you watch fashion videos.
[0325] Example 3: Reducing monitoring costs
[0326] 1. The server verifies the summary of the posted video for inappropriate content.
[0327] 2. If the server determines that there is no problem, the video will be automatically made public and will no longer be subject to manual monitoring.
[0328] Note
[0329] This system utilizes an emotion engine to provide content tailored to user characteristics, optimize advertising revenue, and reduce manual monitoring costs. In addition, by taking into account the user's emotional state, an improved user experience is expected.
[0330] The processing flow will be explained below.
[0331] Step 1:
[0332] A user uses a terminal to upload a video file and its meta information (title, description, etc.) to a server.
[0333] Step 2:
[0334] The server checks the format of the received video file and converts it to a standard format if necessary, as well as checking the video quality and generating metadata (file size, duration, etc.).
[0335] Step 3:
[0336] The server begins analyzing the video content using AI technology, specifically by analyzing the frames of the video and identifying key scenes and important objects.
[0337] Step 4:
[0338] The server performs audio analysis to identify speakers and extract key phrases, and generates a summary of the video based on the results of this analysis.
[0339] Step 5:
[0340] Based on the generated summary, the server assigns tags and keywords to the video and classifies it into a specific category, thereby categorizing the video.
[0341] Step 6:
[0342] The server selects appropriate videos based on the user's viewing history and interests, and generates a recommendation list for the user. This list is sent to the device.
[0343] Step 7:
[0344] The server uses an emotion engine to analyze the user's emotions. Specifically, it collects data such as the user's facial expressions, voice, and viewing history to evaluate their emotional state.
[0345] Step 8:
[0346] The server updates the recommendation list based on the results of the sentiment analysis. Specifically, it prioritizes recommendations of content that users enjoy watching.
[0347] Step 9:
[0348] The server selects video ads based on user characteristics and sentiment analysis results, selects the most suitable ad, and also sets its meta information and playback timing.
[0349] Step 10:
[0350] The server evaluates the relevance of the video content and inserts advertisements at the appropriate time, and the device displays the advertisements.
[0351] Step 11:
[0352] The server automatically determines what to monitor based on the video summary, and detects inappropriate content (e.g., violence, adult content, etc.) from the summary.
[0353] Step 12:
[0354] The server performs a risk assessment and, if automated monitoring determines there are no issues, the video is made public. If manual monitoring is deemed necessary, the monitoring team is notified.
[0355] Step 13:
[0356] The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[0357] Step 14:
[0358] The server filters content that is inappropriate for the specified age and restricts playback on the device, preventing minors from accessing inappropriate content.
[0359] Example 2
[0360] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0361] In today's video streaming services, it is extremely difficult to efficiently process the large number of videos posted by users and provide relevant recommended videos to users. Furthermore, the task of appropriately filtering inappropriate content and ensuring user safety is a heavy burden. Furthermore, there is a demand for sophisticated recommendations and ad delivery based on user sentiment, and a system that can realize these elements in an integrated manner is needed.
[0362] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for uploading videos from a terminal, a means for checking the quality of the received videos and generating metadata, a means for analyzing the content of the videos and generating summaries, a means for categorizing the videos based on the summaries generated by the server and assigning tags and keywords, a means for adding the videos to a recommendation list based on user characteristics, a means for improving the video recommendation function using an emotion engine, a means for selecting and delivering optimal advertisements based on user characteristics, a means for determining whether automatic or manual monitoring is required based on the content of the videos, and a means for applying an age filter based on the estimated age of the user. This makes it possible to efficiently process videos posted by users, provide highly relevant recommendations, deliver optimized advertisements, and provide a safer usage environment.
[0363] "Terminal" is a general term for electronic devices that users use to upload video files and their meta information to a server, and includes, for example, smartphones and personal computers.
[0364] "Server" refers to a central computer system that receives videos uploaded by users and performs processes such as analysis, conversion, summary generation, recommendation and advertisement delivery.
[0365] "Uploading a video" refers to the process in which a user sends a video file from their own device to a server.
[0366] "Quality confirmation" is the process in which the server checks the quality of the video file received, including the format, resolution, and playback time, and converts it into a standard format.
[0367] "Metadata generation" refers to the process by which the server generates additional information, such as file size, playback time, and resolution, based on the received video file.
[0368] "Video content analysis" is the process in which the server uses AI technology to analyze video frames and audio to identify key scenes and important objects.
[0369] "Summary generation" refers to the process in which the server extracts key scenes and important information based on the analysis of the video content and generates a compact summary.
[0370] "Categorization" is the process of classifying videos into specific categories based on server-generated summaries and assigning tags and keywords.
[0371] A "recommendation list" is a list of videos that the server determines to be optimal based on the user's characteristics and provides them to the user.
[0372] An "emotion engine" refers to software or a system that has the function of analyzing emotions from a user's facial expressions, tone of voice, viewing history, etc.
[0373] "Advertisement selection and delivery" is the process in which the server selects the most appropriate advertisement based on user characteristics and sentiment analysis results, and delivers it to the user within the video at the appropriate time.
[0374] "Automatic monitoring" is a function in which the server analyzes the video content and automatically detects and evaluates inappropriate content based on predetermined criteria.
[0375] "Human monitoring" is a process in which a team of human monitors conducts additional verification on videos that are deemed high risk by automated monitoring.
[0376] "Age filter" refers to a feature that restricts content that is displayed or played based on the estimated age of the user.
[0377] This invention is a system that efficiently processes videos posted by users and delivers highly relevant recommendations and optimized advertisements. This system is mainly built based on communications between terminals, a server, and users, and specific embodiments are described below.
[0378] Video Ingestion and Analysis
[0379] 1. The device sends the video uploaded by the user to the server.
[0380] Users send video files and their meta information (title, description, etc.) to the server from their devices such as smartphones or PCs.
[0381] 2. The server receives the video, checks the quality, and generates metadata.
[0382] The server checks the format of the video file, converts it to a standard format if necessary, checks the video quality, and generates metadata (file size, duration, etc.).
[0383] 3. The server uses AI technology to analyze the video content and generate a summary.
[0384] The server analyzes the video frames to identify key scenes and important objects, and performs audio analysis to identify speakers and extract important phrases. Based on this information, a video summary is generated.
[0385] Categorizing the video
[0386] 1. The server categorizes the video based on the generated summary.
[0387] The server classifies videos into specific categories based on tags and keywords extracted from the summaries, and also assigns these tags and keywords to the video data.
[0388] 2. The server creates a recommendation list appropriate for each user.
[0389] The server selects the most suitable videos based on the user's viewing history and interests, and generates a recommendation list, which is then sent to the user's device.
[0390] Emotion analysis and recommendations using an emotion engine
[0391] 1. The server analyzes the user's emotions using an emotion engine.
[0392] The server analyzes the user's emotions based on facial expressions, tone of voice, viewing history, etc.
[0393] 2. The server reflects the results of the sentiment analysis in its video recommendations.
[0394] The server takes the emotional data into account and updates the recommendation list to provide more of the content the user enjoys.
[0395] Video Ad Optimization
[0396] 1. The server selects video ads based on user characteristics.
[0397] The server selects the most suitable advertisement based on the user's age, gender, interests, and emotional analysis results, and also sets the advertisement's meta information and playback timing.
[0398] 2. The server delivers ads within the video.
[0399] The server evaluates the relevance of the video content and inserts ads at the appropriate time, so that relevant ads are displayed when the user watches the video.
[0400] Determining what to monitor
[0401] 1. The server determines whether to monitor the video based on the summary.
[0402] The server detects whether the video summary contains inappropriate content (e.g., violence, adult content, etc.), performs a risk assessment, and determines whether manual supervision is necessary.
[0403] 2. The server will tag the data for manual monitoring as needed.
[0404] Videos that are deemed safe by automatic server monitoring will not be subject to manual monitoring. If manual monitoring is deemed necessary, the monitoring team will be notified.
[0405] Applying an age filter
[0406] 1. The server applies an age filter based on the user's estimated age.
[0407] The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[0408] 2. The server filters inappropriate content.
[0409] The server prevents content that is not appropriate for the specified age from being played on the user's device, thereby preventing minors from accessing inappropriate content.
[0410] Specific examples
[0411] Example 1: User-Submitted Pet Videos
[0412] 1. The user (device) uploads a video of their dog playing at home.
[0413] 2. The server receives the video, performs quality checks, and generates metadata.
[0414] 3. The server analyzes the video and generates keywords such as "pet," "dog," and "play."
[0415] 4. The server uses these keywords to recommend relevant pet videos to the user.
[0416] 5. The server performs sentiment analysis on the user, detects whether the user is happy, and recommends more pet videos based on that.
[0417] Example 2: Fashion-related video ads
[0418] 1. The server identifies that the user has fashion-related interests.
[0419] 2. The server uses an emotion engine to confirm that the user has positive emotions toward the fashion content.
[0420] 3. The server selects fashion-related video ads and inserts them into the video at the appropriate time.
[0421] 4. The device displays relevant ads while you watch fashion videos.
[0422] Example 3: Reducing monitoring costs
[0423] 1. The server verifies the summary of the posted video for inappropriate content.
[0424] 2. If the server determines that there is no problem, the video will be automatically made public and will no longer be subject to manual monitoring.
[0425] Prompt Sentence Examples
[0426] 1. Analyze the video content and extract key scenes and important phrases.
[0427] 2. Generate a recommendation list based on the user's viewing history.
[0428] 3. Use a sentiment engine to analyze user sentiment and provide appropriate content.
[0429] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0430] Step 1:
[0431] The device sends the video file and its meta information (title, description, etc.) to the server.
[0432] Input: A user uploads a video file and its meta information from their device.
[0433] Processing: The device sends the video file to the server.
[0434] Output: The video file and associated meta information are saved on the server.
[0435] Step 2:
[0436] The server checks the format of the video file received and converts it to a standard format if necessary.
[0437] Input: The submitted video file.
[0438] Processing: The server analyzes the format of the video file and converts it to a standard format (e.g., adjusting the resolution, converting the file format).
[0439] Output: Video files converted to standard format.
[0440] Step 3:
[0441] The server checks the quality of the received video and generates metadata.
[0442] Input: Video files converted to standard formats.
[0443] Processing: The server checks the quality of the video (e.g., resolution, duration, file size, etc.) and generates metadata (e.g., resolution, duration, file size, frame rate).
[0444] Output: Quality-checked video files and their metadata.
[0445] Step 4:
[0446] The server uses AI technology to analyze the video content and generate a summary.
[0447] Input: Quality-checked video files and metadata.
[0448] Processing: The server uses generative AI models to analyze video frames and audio, extract key scenes and phrases, identify speakers in the video, and generate a summary.
[0449] Output: Summary of video content and analysis results.
[0450] Step 5:
[0451] The server categorizes the video based on the generated summary and assigns tags and keywords.
[0452] Input: Summary of video content and analysis results.
[0453] Processing: Based on the summary, the server assigns tags and keywords (e.g., "pets," "dogs," "play," etc.) to the video and classifies it into a specific category.
[0454] Output: Video data with tags and keywords.
[0455] Step 6:
[0456] The server creates a recommendation list suitable for each user.
[0457] Input: Video data with tags and keywords, user viewing history.
[0458] Processing: The server selects relevant videos based on the user's viewing history and interests and generates a recommendation list.
[0459] Output: A list of recommendations suitable for the user.
[0460] Step 7:
[0461] The server analyzes the user's emotions using an emotion engine.
[0462] Input: User's viewing history, facial expressions, and tone of voice.
[0463] Processing: The server uses an emotion engine to analyze the user's facial expressions and tone of voice and extract emotional data.
[0464] Output: User sentiment analysis results.
[0465] Step 8:
[0466] The server reflects the results of the sentiment analysis in its video recommendations.
[0467] Input: User sentiment analysis results, recommendation list.
[0468] Processing: The server takes into account the emotion data and updates the recommendation list to provide more content that the user enjoys.
[0469] Output: An updated recommendation list reflecting the sentiment data.
[0470] Step 9:
[0471] The server selects video advertisements based on user characteristics.
[0472] Input: User's age, gender, interests, and sentiment analysis results.
[0473] Processing: The server selects the most suitable advertisement based on the user's characteristics and sets the advertisement's meta information and playback timing.
[0474] Output: Advertising data based on user characteristics.
[0475] Step 10:
[0476] The server delivers ads within the video.
[0477] Input: Advertising data based on user characteristics, video being watched.
[0478] Processing: The server evaluates the relevance of the video content and inserts advertisements at the appropriate time.
[0479] Output: Ads inserted during viewing.
[0480] Step 11:
[0481] The server determines whether the video should be monitored based on the summary.
[0482] Input: A summary of the video content.
[0483] Processing: The server detects inappropriate content (e.g., violence, adult content) from the video summary and performs a risk assessment.
[0484] Output: The result of automated or manual monitoring.
[0485] Step 12:
[0486] The server will tag the monitors manually as needed.
[0487] Input: Verdict of automated or manual monitoring.
[0488] Processing: The server automatically monitors and publishes videos that it determines to be problem-free, and notifies the monitoring team if it determines that there is a problem.
[0489] Output: published video or notification to surveillance team.
[0490] Step 13:
[0491] The server applies an age filter based on the user's estimated age.
[0492] Input: User registration information, activity history, video summary.
[0493] Processing: The server estimates the user's age and detects age-restricted keywords (e.g., "violence" and "adult") from the video summary.
[0494] Output: The filtered content.
[0495] Step 14:
[0496] The server filters inappropriate content.
[0497] Input: Content after age filter applied.
[0498] Processing: The server filters content that is inappropriate for the specified age group and prevents it from being played on the user's device.
[0499] Output: Content that is restricted from being accessed by minors.
[0500] (Application example 2)
[0501] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0502] In modern content distribution services, the large number of users posting videos necessitates efficient video summarization and the creation of recommendation lists tailored to each user. It is also important to optimize ad delivery, reduce monitoring costs, and provide content based on user characteristics. However, no system exists that can simultaneously address all of these challenges. In particular, video recommendations and ad delivery that take user emotions into account have yet to be realized, so an effective solution is needed.
[0503] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0504] In this invention, the server includes: a means for uploading videos from a terminal; a means for checking the quality of the received videos and generating metadata; a means for analyzing the video content and generating a summary; a means for categorizing the videos based on the generated summary and assigning tags and keywords; a means for adding the videos to a recommendation list based on user characteristics; a means for selecting and assigning optimal advertisements based on user characteristics; a means for determining whether automatic or manual monitoring is required based on the video content; a means for applying an age filter based on the estimated age of the user; a means for analyzing the user's emotions using emotion recognition technology and updating the video recommendation list based on the analysis; and a means for inserting video advertisements related to specific video categories into videos at appropriate times. This enables efficient creation of video summaries and recommendation lists, optimization of advertisement distribution, reduction of monitoring costs, and provision of advanced content that takes user emotions into consideration.
[0505] A "terminal" is an electronic device such as a smartphone or computer that a user uses to upload videos.
[0506] "Server" refers to a computer system on a network that receives videos, analyzes them, generates summaries, categorizes them, creates recommendation lists, selects advertisements, determines whether they are monitored automatically or manually, applies age filters, analyzes emotions, and delivers advertisements.
[0507] "Quality check" is a process in which the server checks the format, image quality, sound quality, etc. of the uploaded video and converts it to a standard format if necessary.
[0508] "Metadata generation" is the process by which the server generates the video's file size, playback time, and other related information.
[0509] "Summary generation" is a process in which the server analyzes the content of a video and creates a short summary of the video based on key scenes, important objects, and audio.
[0510] "Categorization" is a process in which the server classifies videos into specific categories based on the summary information generated and assigns related tags and keywords.
[0511] A "recommendation list" is a list in which the server selects specific videos based on the user's viewing history and interests and provides them to the user.
[0512] "Advertisement selection" is a process in which the server selects the most suitable advertisement based on the user's age, gender, interests, and emotional analysis results, and sets its meta information and playback timing.
[0513] "Automatic monitoring" is a process in which the server detects whether a video contains inappropriate content based on its summary information and determines whether manual monitoring is necessary.
[0514] "Age filter" is a process in which the server controls the viewing of inappropriate content based on the estimated age of the user.
[0515] "Emotion recognition technology" is a technology that allows the server to analyze the user's facial expressions, tone of voice, etc. to understand the user's emotional state.
[0516] A "video category" is a specific genre or theme that the server uses to categorize videos, such as pets, fashion, cooking, etc.
[0517] The system for implementing this invention is configured by linking user terminals, a server, and a network, and enables efficient video summarization, improved recommendation functionality, optimized ad distribution, reduced monitoring costs, and a safe usage environment.
[0518] First, a user uploads a video from a device such as a smartphone or PC. This device is used to send the video taken by the user to the server.
[0519] The server temporarily stores the received video and checks its quality. Specifically, it checks the video format and converts it to a standard format if necessary. It also checks the video's image quality and sound quality and generates metadata (such as file size and playback time). This process uses video editing libraries (e.g., OpenCV) and audio analysis libraries.
[0520] The server then uses AI technology to analyze the video content. This analysis involves identifying important scenes in each frame of the video and extracting key phrases from the audio data. The AI technology used here utilizes computer vision for video analysis and a speech recognition engine (e.g., Google Speech-to-Text API) for audio analysis. The generated summary is then compiled in a format that is easy for the user to view.
[0521] The server then categorizes the video based on the generated summary and assigns relevant tags and keywords to it, using natural language processing (NLP) techniques based on information extracted from the summary.
[0522] The server then generates a recommendation list based on the user's characteristics. The most suitable videos are selected based on the user's viewing history and interests, and this list is sent to the user's device. The recommendation engine uses machine learning models to analyze user behavior data.
[0523] Furthermore, the server uses emotion recognition technology to analyze the user's emotions. This analysis involves analyzing the user's facial expressions and tone of voice. The emotion engine uses, for example, the FER (Facial Emotion Recognition) library. Based on the analysis results, the video recommendation list according to the emotion is updated.
[0524] The server selects the most suitable advertisement based on the user's characteristics and inserts it into the video at the appropriate time. Ad selection takes into account age, gender, interests, and emotional analysis results. Ad delivery also uses an algorithm that evaluates the relevance of the advertisement to the video content.
[0525] Finally, the server has a means to determine whether to monitor automatically or manually based on the video content. The automatic monitoring system analyzes the summary of the uploaded video to detect inappropriate content, such as illegal content. If necessary, it notifies the manual monitoring team.
[0526] A concrete example is the automatic recommendation of pet videos. When a user uploads a video of their dog playing at home, the system receives and analyzes the video, generating tags such as "pet," "dog," and "play." Based on the user's past viewing history and sentiment analysis data, the system recommends new pet videos.
[0527] Example of an input prompt for a generative AI model:
[0528] "Enter the video title, description, and uploaded file path. Example: Title: 'Dogs playing at home', Description: 'Dogs playing filmed at home', File path: 'path / to / video.mp4'"
[0529] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0530] Step 1:
[0531] Videos are uploaded from the device
[0532] A user shoots a video using a smartphone or computer and accesses the upload page. The prompt message is "Please enter the video title, description, and uploaded file path. Example: Title: 'Dog playing at home', Description: 'Dog playing filmed at home', File path: 'path / to / video.mp4'." When the user enters the title, description, and video file and presses the upload button, the video file and metadata are sent to the server. The input is the video file and its metadata specified by the user, and the output is the video file and its metadata stored on the server.
[0533] Step 2:
[0534] The server checks the quality of the received video and generates metadata.
[0535] The server temporarily stores the received video file and checks its format, image quality, and sound quality. For example, it uses OpenCV to check the video resolution and pixel quality and converts it to a standard format if necessary. It also uses an audio analysis library to check the sound quality. Metadata such as file size, playback time, and resolution are generated. The input is the received video file, and the output is the generated metadata.
[0536] Step 3:
[0537] The server analyzes the video content and generates a summary
[0538] The server uses AI technology to analyze the video content. Each frame of the video is analyzed using computer vision technology to identify key scenes. For example, OpenCV and deep learning models are used to detect key objects and movements. The audio is converted to text using the Google Speech-to-Text API, and key phrases are extracted. A summary is generated based on this information. The input is each frame of the video and audio data, and the output is a video summary.
[0539] Step 4:
[0540] The server categorizes the video based on the generated summary and assigns tags and keywords.
[0541] The server classifies videos into specific categories based on information obtained from the video summaries and assigns associated tags and keywords. For example, natural language processing techniques are used to extract keywords obtained from the summaries and map them to predefined categories. The input is the generated video summaries, and the output is the classified video categories and assigned tags and keywords.
[0542] Step 5:
[0543] The server adds videos to the recommendation list based on the user's characteristics.
[0544] The server analyzes the user's viewing history, interests, and past viewing data to select the most suitable videos and generate a recommendation list. Here, a machine learning model analyzes the user's behavioral data and lists highly relevant videos. For example, a collaborative filtering algorithm is used. The input is the user's viewing history and interest data, and the output is a recommendation list.
[0545] Step 6:
[0546] The server uses emotion recognition technology to analyze the user's emotions and updates the video recommendation list based on that.
[0547] The server performs facial expression analysis and voice tone analysis to recognize the user's emotions. For example, it uses the FER (Facial Emotion Recognition) library to analyze facial expressions in frames and identify the user's emotions. Based on the results, it adds more types of videos that the user enjoys to the recommendation list. The input is each frame of the video and audio data, and the output is an updated recommendation list.
[0548] Step 7:
[0549] The server selects the most suitable advertisement based on the user's characteristics and inserts it into the video at the appropriate time.
[0550] The server selects the most suitable advertisement based on the user's age, gender, interests, and emotional analysis results, and determines the timing of its insertion into the video. For example, collaborative filtering and emotional analysis data can be combined to identify the user's areas of interest, select highly relevant advertisements, and incorporate them into the video frame. The input is the user's characteristics and the results of the emotional analysis, and the output is a video containing the inserted advertisement.
[0551] Step 8:
[0552] The server determines whether to monitor automatically or manually based on the video content.
[0553] The server uses the generated summary to check for inappropriate content in the video using an automated monitoring system. For example, it uses content analysis algorithms to detect problematic scenes or audio, performs a risk assessment, and notifies a human monitoring team if necessary. The input is the generated video summary, and the output is the monitoring decision.
[0554] Step 9:
[0555] The server applies an age filter based on the user's estimated age.
[0556] The server estimates the user's age based on their registration information and activity history, and restricts access to inappropriate content. For example, it extracts age-related keywords from the summary information and filters videos that do not match the user's age. The input is the user's estimated age and the video summary, and the output is a filtered list of videos.
[0557] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0558] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0559] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0560] [Second embodiment]
[0561] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0562] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0563] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0564] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0565] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0566] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0567] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0568] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0569] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0570] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0571] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0572] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0573] This invention is a system that efficiently summarizes videos posted by users and, based on the summaries, enables recommendation functions, advertisement distribution, reduced monitoring costs, and a safe usage environment. This system is built on communications between terminals, servers, and users. Below, the processing of the system's program is explained in natural language, with specific examples.
[0574] Video Ingestion and Analysis
[0575] 1. The device sends the video uploaded by the user to the server.
[0576] Device: The user uploads the video file and its meta information (title, description, etc.) to the server from a device such as a smartphone or PC.
[0577] 2. The server receives the video and performs the ingest process.
[0578] Server: Checks the format of the video file and converts it to a standard format if necessary. It also checks the video quality and generates metadata (file size, duration, etc.).
[0579] 3. The server uses AI technology to analyze the video content and generate a summary.
[0580] Server: Analyzes video frames to identify key scenes and important objects. Also, analyzes audio to identify speakers and extract important phrases. Then, it generates a video summary based on this information.
[0581] Categorizing the video
[0582] 1. The server categorizes the video based on the generated summary.
[0583] Server: Classifies videos into specific categories based on tags and keywords extracted from the summaries, and assigns these tags and keywords to the video data.
[0584] 2. The server creates a recommendation list suitable for each user.
[0585] Server: Based on the user's viewing history and interests, the server selects the most suitable videos and generates a recommendation list. This list is sent to the user's device.
[0586] Video Ad Optimization
[0587] 1. The server selects video ads based on user characteristics.
[0588] Server: Selects the most suitable advertisement based on the user's characteristics such as age, gender, and interests. Also sets the advertisement meta information and playback timing.
[0589] 2. The server delivers ads within the video.
[0590] Server: Evaluates the relevance of the video content and inserts ads at the appropriate time. When a user watches a video, relevant ads are displayed.
[0591] Determining what to monitor
[0592] 1. The server determines whether to monitor the video based on the summary.
[0593] Server: Detects whether the video summary contains inappropriate content (e.g. violence, adult content, etc.), performs a risk assessment, and determines whether manual supervision is required.
[0594] 2. The server will tag the data for manual monitoring as needed.
[0595] Server: Videos that are deemed safe by automatic monitoring are exempt from manual monitoring. If manual monitoring is deemed necessary, the monitoring team is notified.
[0596] Applying an age filter
[0597] 1. The server applies an age filter based on the user's estimated age.
[0598] Server: Estimates the user's age from their registration information and activity history. Detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[0599] 2. The server filters inappropriate content.
[0600] Server: Content that is not appropriate for the specified age group will not be played on the user's device, thereby preventing minors from accessing inappropriate content.
[0601] Specific examples
[0602] Example 1: User-Submitted Pet Videos
[0603] 1. The user (device) uploads a video of their dog playing at home.
[0604] 2. The server receives the video, performs quality checks, and generates metadata.
[0605] 3. The server analyzes the video and generates keywords such as "pet," "dog," and "play."
[0606] 4. The server uses these keywords to recommend pet videos relevant to the user.
[0607] Example 2: Fashion-related video ads
[0608] 1. The server identifies that the user has fashion-related interests.
[0609] 2. The server selects fashion-related video ads and inserts them into the video at the appropriate time.
[0610] 3. The device displays relevant ads while you watch fashion videos.
[0611] Example 3: Reducing monitoring costs
[0612] 1. The server verifies the summary of the posted video for inappropriate content.
[0613] 2. If the server determines that there is no problem, the video will be automatically made public and will no longer be subject to manual monitoring.
[0614] This system enables the provision of content tailored to user characteristics, optimization of advertising revenue, and reduction of human monitoring costs.
[0615] The processing flow will be explained below.
[0616] Step 1:
[0617] A user uses a terminal to upload a video file and its meta information (title, description, etc.) to a server.
[0618] Step 2:
[0619] The server checks the format of the video file received and converts it to a standard format if necessary, checks the video quality, and generates metadata (file size, playback time, etc.).
[0620] Step 3:
[0621] The server begins analyzing the video content using AI technology, specifically by analyzing the frames of the video and identifying key scenes and important objects.
[0622] Step 4:
[0623] The server performs audio analysis to identify speakers and extract key phrases, and generates a video summary based on the results of this analysis.
[0624] Step 5:
[0625] Based on the generated summary, the server assigns tags and keywords to the video and classifies it into a specific category, thereby categorizing the video.
[0626] Step 6:
[0627] The server selects appropriate videos based on the user's viewing history and interests, and generates a recommendation list for the user. This list is sent to the device.
[0628] Step 7:
[0629] The server selects the most suitable video ad based on user characteristics (age, gender, interests, etc.) and also sets the ad's meta information and playback timing.
[0630] Step 8:
[0631] The server evaluates the relevance of the video content and inserts ads at the appropriate time, and displays relevant ads when the device is watching the video.
[0632] Step 9:
[0633] Based on the summary, the server determines whether the video is suitable for automatic monitoring or requires manual monitoring. Inappropriate content (e.g., violence, adult content, etc.) is detected.
[0634] Step 10:
[0635] The server performs a risk assessment and, if it determines that there are no problems with automated monitoring, it makes the video public. If it determines that manual monitoring is necessary, it notifies the monitoring team.
[0636] Step 11:
[0637] The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the summary.
[0638] Step 12:
[0639] The server filters content that is inappropriate for the specified age and prevents the video from being played on the device, thereby preventing minors from accessing inappropriate content.
[0640] Example 1
[0641] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0642] Modern video streaming services require efficient analysis and classification of the vast amount of videos uploaded by users, and the provision of appropriate recommendations and advertisements. However, manual video analysis and recommendation generation requires a great deal of effort and time, resulting in high operational costs, making efficient and automated processing methods necessary. Automation is also required for providing personalized advertisements based on user characteristics, filtering content suitable for underage users, and monitoring video content. An efficient and reliable system is needed to solve these challenges.
[0643] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0644] In this invention, the server includes a means for uploading videos from a terminal, a means for converting the format of the received video, checking the quality, and generating metadata, a means for analyzing the content of the video using a generation AI model and generating a summary, a means for categorizing the video based on the generated summary and assigning tags and keywords, a means for adding the video to a recommendation list based on user characteristics, a means for selecting and delivering optimal advertisements based on user characteristics, a means for determining whether automatic or manual monitoring is required based on the video content, and a means for applying an age filter based on the estimated age of the user. This enables efficient analysis of videos uploaded by users, provision of appropriate recommendations and advertisements, automated content monitoring, and age-appropriate filtering.
[0645] A "terminal" is a device used by a user, such as a smartphone or computer, used to upload videos.
[0646] The "server" is a central device that receives, analyzes, and processes video data over the network, and plays a central role in this system.
[0647] A "video" is a file containing a series of video and audio files uploaded by a user.
[0648] "Format conversion" is the process of converting received video files into a standard format.
[0649] "Quality check" is the process of checking the quality of the received video, such as its resolution and playback time.
[0650] "Metadata" refers to various information related to a video file (e.g., file size, playback time, resolution, etc.).
[0651] A "generative AI model" is an artificial intelligence technology used to analyze video content and generate summaries.
[0652] "Video analysis" is the process of analyzing video frames and audio to identify important scenes and objects.
[0653] "Summary" refers to a short summary of the main content extracted from the results of video analysis.
[0654] "Tags" refer to keywords or labels that are added to indicate the content of a video.
[0655] "Keywords" refer to words or short phrases related to the video content.
[0656] A "recommendation list" is a list of videos selected based on the user's characteristics and interests.
[0657] "Advertisement" refers to information including product promotions and service information, which is inserted into videos so that users can view them.
[0658] "Automatic monitoring" is the process of automatically checking video content to detect inappropriate content.
[0659] "Manual monitoring" is a manual monitoring process carried out by specialized staff when automatic monitoring is difficult to judge.
[0660] "Estimated age" is the age calculated based on the user's registration information and activity history.
[0661] An "age filter" is a mechanism that provides only appropriate content based on the user's estimated age.
[0662] This invention is a system that efficiently summarizes videos posted by users and provides recommendation functions, advertisement distribution, reduced monitoring costs, and a safe usage environment based on the summaries. This system is built on communications between terminals, a server, and users.
[0663] Video Ingestion and Analysis
[0664] When a user uploads a video to a server from a device, the user uses a device such as a smartphone or PC. Specifically, the user sends the video file and its meta information (title, description, etc.) to the server through an interface provided by the device.
[0665] The server receives video files uploaded by users. The received video files are first converted into a different format. Software such as ffmpeg is used for this process. Next, the quality of the video is checked and metadata (file size, resolution, playback time, etc.) is generated based on that quality.
[0666] The server uses a generative AI model to analyze the video content and generate a summary. Specifically, it uses an object detection model such as YOLO (You Only Look Once) to identify key scenes and important objects, and then performs audio analysis using the Google Cloud Speech-to-Text API. Then, based on this information, it uses a natural language generation model (e.g., GPT-3) to generate a video summary.
[0667] Video categorization and recommendations
[0668] The server categorizes the videos based on the generated summaries. The videos are classified into specific categories based on tags and keywords extracted from the summaries. The classified videos are then assigned tags and keywords to improve searchability and filtering capabilities.
[0669] The server then generates a recommendation list based on the user's characteristics. Taking into account the user's viewing history and interests, the server selects the most suitable videos to create the recommendation list. This recommendation list is then sent to the user's device, where the user can watch the recommended videos through an interface.
[0670] Ad optimization and delivery
[0671] The server selects the most suitable advertisement based on the video summary and user characteristics. Specifically, it selects advertisements based on the user's characteristics such as age, gender, and interests, and sets the meta information and playback timing. The server references a database provided by the advertiser to select advertisements.
[0672] The server delivers advertisements at appropriate times within the video. For example, while watching a video about pets, it displays an advertisement for related pet food. This can be expected to have a high advertising effect.
[0673] Determining who should be monitored and applying age filters
[0674] The server determines whether a video should be subject to automated or manual monitoring based on the summary. The video summary is used to detect whether it contains inappropriate content and perform a risk assessment. Based on this, the process of determining whether to subject the video to automated or manual monitoring uses natural language processing technology.
[0675] Furthermore, the server applies an age filter based on the user's estimated age. The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the video summary. Inappropriate content is filtered so that it cannot be played on devices belonging to underage users. Based on these filtering rules, only appropriate videos are provided.
[0676] Specific examples
[0677] Example 1: User-Submitted Pet Videos
[0678] A user uploads a video of their dog playing at home. The server receives the video, checks the quality, and generates metadata. The server then analyzes the video and generates keywords such as "pet," "dog," and "play." The server then recommends relevant pet videos to the user based on these keywords.
[0679] Example 2: Fashion-related video ads
[0680] The server identifies the user's interest in fashion. The server then selects fashion-related video ads and inserts them into the video at the appropriate time. The relevant ads are then displayed on the user's device while they are watching the fashion video.
[0681] Example 3: Reducing monitoring costs
[0682] The server verifies the summary of the posted video to see if it contains inappropriate content. If it determines there are no problems, it automatically publishes the video and removes it from the scope of human monitoring.
[0683] Prompt Sentence Examples
[0684] "Recommend other videos related to this pet video."
[0685] "Insert fashion-related ads at the best possible time."
[0686] Please determine if this video is inappropriate.
[0687] This system makes it possible to provide content tailored to user characteristics, optimize advertising revenue, and reduce manual monitoring costs.
[0688] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0689] Step 1:
[0690] The device sends the video uploaded by the user to the server. The user clicks the upload button through the device interface, selects the video file (e.g., mp4, avi) and its meta information (title, description), and sends it to the server.
[0691] Input: Video file and meta information
[0692] Output: Video file and meta information sent to the server
[0693] Step 2:
[0694] The server receives the video, converts the format, checks the quality, and generates metadata. The server checks the format of the received video file and converts it to a standard format (e.g., mp4). This process uses tools such as ffmpeg. It also checks the quality of the video, such as resolution, playback time, and file size, and generates this information as metadata.
[0695] Input: Video file and meta information sent to the server
[0696] Output: Format converted video file and generated metadata
[0697] Step 3:
[0698] The server uses a generative AI model to analyze the video content and generate a summary. Specifically, it uses an object detection model such as YOLO to analyze video frames and identify key scenes and important objects. It also performs audio analysis using the Google Cloud Speech-to-Text API to identify speakers and extract key phrases. It then uses this data to generate a video summary using a natural language generation model (e.g., GPT-3).
[0699] Input: Format converted video file and metadata
[0700] Output: Video summary
[0701] Step 4:
[0702] The server categorizes the video based on the generated summary and assigns tags and keywords. Based on the information obtained from the summary, the server classifies the video into an appropriate category (e.g., pets, sports, entertainment), and assigns tags and keywords related to that category to the video data.
[0703] Input: Video summary
[0704] Output: Categorized videos with tags and keywords
[0705] Step 5:
[0706] The server adds videos to a recommendation list based on user characteristics. The server analyzes the user's viewing history and interests, selects the most suitable videos, and generates a recommendation list. This list is sent to the user's device.
[0707] Input: User's viewing history and interests, as well as categorized videos, tags, and keywords
[0708] Output: Recommendation list based on user characteristics
[0709] Step 6:
[0710] The server selects and delivers the most suitable advertisement based on the user's characteristics. The server selects the appropriate advertisement based on the user's characteristics information such as age, gender, and interests. Specifically, it retrieves the advertisement that best suits the user's characteristics from the advertisement database and sets its meta information and playback timing.
[0711] Input: User characteristics (age, gender, interests, etc.) and advertising database
[0712] Output: Optimal advertisement based on user characteristics
[0713] Step 7:
[0714] The server determines whether the video should be monitored automatically or manually based on its content. The server then checks the video summary for inappropriate content and assesses the content risk. This process uses natural language processing technology.
[0715] Input: Video summary
[0716] Output: Judgment results of automatic or manual monitoring
[0717] Step 8:
[0718] The server applies an age filter based on the user's estimated age. The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords from the video summary. Inappropriate content is filtered so that it cannot be played on devices belonging to minors.
[0719] Input: User registration information, activity history, video summary
[0720] Output: Age-filtered video list
[0721] (Application example 1)
[0722] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0723] In recent years, with the proliferation of video content, a huge amount of videos has been uploaded to online platforms. However, there is a lack of systems that can efficiently summarize videos and recommend appropriate content to users. Furthermore, manually monitoring video content and filtering inappropriate content is extremely time-consuming and costly. Furthermore, optimization of advertisements based on user characteristics is insufficient, and efficient advertisement delivery is required. To address these issues, a system is needed that automates video analysis and recommendation, advertisement optimization, and content monitoring, while providing a safe viewing environment.
[0724] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0725] In this invention, the server includes a means for uploading videos from a terminal, a means for checking the quality of the received videos and generating metadata, a means for analyzing the video content and generating a summary, a means for categorizing the videos based on the generated summary and assigning tags and keywords, a means for adding the videos to a recommendation list based on user characteristics, a means for selecting and assigning optimal advertisements based on user characteristics, a means for determining whether automatic or manual monitoring is required based on the video content, a means for applying an age filter based on the estimated age of the user, a means for saving extracted frame images and transcribing and analyzing audio, a means for providing a function for recommending similar content based on responses from the server, and a means for inserting advertisements and optimizing based on user characteristics as part of a content distribution service. This enables efficient video summarization and recommendation functions, optimal advertisement distribution, reduced monitoring costs, and a safe usage environment.
[0726] - "Device" refers to the device to which the video is uploaded, such as a smartphone or computer.
[0727] "Server" refers to the computing system that processes the received video, checks the quality, generates metadata, and analyzes the video content.
[0728] "Video content analysis" refers to the process of analyzing video frames and audio data to extract summaries, important scenes, and objects.
[0729] "Generating a summary" refers to extracting the main points from the analyzed video data and summarizing them concisely.
[0730] "Categorizing videos" refers to classifying videos into specific categories based on the generated summary and assigning relevant tags and keywords.
[0731] "Add to recommendation list" refers to selecting suitable videos based on the user's viewing history and interests and adding them to the display list.
[0732] "Selecting and delivering ads" refers to selecting the most suitable ads based on user characteristics and inserting them while the video is playing.
[0733] "Determine whether to use automated or manual monitoring" refers to determining whether a video contains inappropriate content based on a summary of the video and deciding whether to use automated or, if necessary, manual monitoring.
[0734] "Applying age filters" refers to restricting the content available to a user based on their estimated age.
[0735] "Save frame images" means saving still images extracted from a video, usually at 1-second intervals.
[0736] "Audio transcription and analysis" refers to the process of converting video audio into text data and extracting key phrases and keywords.
[0737] Providing a "similar content recommendation function" refers to a function that automatically suggests relevant videos based on a user's viewing history and characteristics.
[0738] "Ad insertion and optimization based on user characteristics" refers to inserting ads into videos at the appropriate time and providing the optimal advertising plan based on the user's characteristics.
[0739] This invention provides a system that efficiently summarizes videos posted by users and, based on the summaries, enables recommendation functions, advertisement distribution, reduced monitoring costs, and a safe usage environment. Detailed modes for carrying out the invention are described below.
[0740] Uploading videos from your device
[0741] Users upload video files to the server using devices such as smartphones or PCs. At this time, the video file and its meta information (title, description, etc.) are also sent.
[0742] Server receives and processes video
[0743] The server checks the quality of the received video files, converts them to a standard format if necessary, and generates metadata (file size, playback time, etc.) for the video.
[0744] Video content analysis and summary generation
[0745] The server analyzes the video frames to identify key scenes and important objects, and performs audio analysis to identify speakers and extract key phrases. Based on this information, a video summary is generated. This process includes saving the captured frame images and transcribing and analyzing the audio extracted from the video.
[0746] Video categorization and recommendations
[0747] The server categorizes the videos based on the generated summaries and assigns tags and keywords. Based on the user's characteristics, it selects the most suitable videos and generates a recommendation list. Based on the response from the server, a recommendation function for similar content is provided.
[0748] Ad optimization and insertion
[0749] The server selects the most suitable advertisement based on the user's characteristics and inserts it at a timing relevant to the video content. As part of the content distribution service, it performs advertisement insertion and optimization based on user characteristics.
[0750] Video monitoring and age filtering
[0751] The server determines whether to monitor the content based on the video summary, automatically or manually, and applies an age filter based on the user's estimated age to filter out inappropriate content.
[0752] Hardware and software used
[0753] The following hardware and software are used to implement this system.
[0754] Hardware: Smartphones, PCs
[0755] Software: MoviePy (video editing library), Vosk (voice recognition library), OpenCV (image processing library)
[0756] Specific examples
[0757] Example 1: User-Submitted Pet Videos
[0758] When a user uploads a video of their dog playing at home, the server receives the video, checks the quality, and generates metadata. It then analyzes the video to generate keywords such as "pet," "dog," and "play," and recommends related pet videos.
[0759] Example 2: Fashion-related video ads
[0760] The server identifies that the user is interested in fashion, selects fashion-related video advertisements, and inserts them into the videos at the appropriate time. The relevant advertisements are displayed while the user is watching the fashion video.
[0761] Prompt Sentence Examples
[0762] text
[0763] Analyzes the user-uploaded video "example_video.mp4" and generates a summary. Then, recommends similar videos and inserts relevant ads. Finally, applies supervision and age filters.
[0764] In this way, the present invention makes it possible to provide efficient video summarization and recommendation functions, optimal advertisement distribution, reduced monitoring costs, and a safe usage environment.
[0765] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0766] Step 1:
[0767] A user uses a terminal to upload a video file to a server.
[0768] Input: Video file, meta information (title, description, etc.)
[0769] Output: Video files and meta information are saved on the server
[0770] Specific operation: A user selects a video file and presses the upload button on their smartphone or PC. The device then sends the selected video file and its meta information to the server as an HTTP request.
[0771] Step 2:
[0772] The server checks the quality of the received video file and generates metadata.
[0773] Input: Uploaded video file, meta information
[0774] Output: Video quality evaluation results, metadata (file size, playing time, etc.)
[0775] What it does: The server checks the format of the video file and converts it to a standard format if necessary. It also evaluates the video quality and generates metadata such as file size, duration, and resolution.
[0776] Step 3:
[0777] The server performs frame analysis and audio analysis of the video content and generates a summary.
[0778] Input: Video file
[0779] Output: Video summary, list of important scenes and objects, audio transcription
[0780] Specific operations: Extract key scenes from each frame of a video using MoviePy and OpenCV. Transcribe audio data using Vosk and extract key phrases and keywords. Combine these to generate a video summary.
[0781] Step 4:
[0782] The server categorizes the video based on the generated summary and assigns tags and keywords.
[0783] Input: Video summary
[0784] Output: Video data with tags and keywords
[0785] How it works: The server classifies videos into specific categories based on phrases and keywords extracted from the summaries, automatically assigning tags such as "pets," "dogs," and "play."
[0786] Step 5:
[0787] The server generates a recommendation list based on user characteristics.
[0788] Input: User viewing history, interests, summary data
[0789] Output: Recommendation list for each user
[0790] How it works: The server analyzes the user's past viewing history and interests (e.g., pet videos) and selects relevant videos from the summary data, thereby generating a customized recommendation list for each user.
[0791] Step 6:
[0792] The server selects the most suitable advertisement based on user characteristics and inserts it at the appropriate time within the video.
[0793] Input: User characteristics, video summary, relevant advertising data
[0794] Output: Video with ads
[0795] Specific operation: The server selects the most suitable advertisement based on the user's characteristics such as age, gender, and interests, and inserts the selected advertisement at the relevant timing in the video (e.g., scene change).
[0796] Step 7:
[0797] The server determines which videos to monitor based on the video summary.
[0798] Input: Video summary
[0799] Output: Judgment result of whether the target is automatic or manual monitoring
[0800] How it works: The server evaluates the summary and determines whether it contains inappropriate content. If the automated monitoring finds no problems, the video is left public. If manual monitoring is deemed necessary, the monitoring team is notified.
[0801] Step 8:
[0802] The server applies an age filter based on the user's estimated age.
[0803] Input: User registration information, activity history, video summary
[0804] Output: Filtered video content
[0805] Specific operation: The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords from the video summary. It then filters the video to prevent underage users from accessing inappropriate content.
[0806] Examples of prompt statements
[0807] Analyzes the user-uploaded video "example_video.mp4" and generates a summary. Then, recommends similar videos and inserts relevant ads. Finally, applies supervision and age filters.
[0808] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0809] This invention is a system that efficiently summarizes videos posted by users, improving recommendation functionality, optimizing ad delivery, reducing monitoring costs, and providing a safe usage environment. Furthermore, by combining it with an emotion engine that analyzes user emotions, more advanced recommendations and ad delivery become possible. This system is built on communications between terminals, servers, and users. Below, we explain the system's program processing in natural language, including specific examples.
[0810] Video Ingestion and Analysis
[0811] 1. The device sends the video uploaded by the user to the server.
[0812] Device: The user uploads the video file and its meta information (title, description, etc.) to the server from a device such as a smartphone or PC.
[0813] 2. The server receives the video and performs the ingest process.
[0814] Server: Checks the format of the video file and converts it to a standard format if necessary. It also checks the video quality and generates metadata (file size, duration, etc.).
[0815] 3. The server uses AI technology to analyze the video content and generate a summary.
[0816] Server: Analyzes video frames to identify key scenes and important objects. Also, analyzes audio to identify speakers and extract important phrases. Then, it generates a video summary based on this information.
[0817] Categorizing the video
[0818] 1. The server categorizes the video based on the generated summary.
[0819] Server: Classifies videos into specific categories based on tags and keywords extracted from the summaries, and assigns these tags and keywords to the video data.
[0820] 2. The server creates a recommendation list suitable for each user.
[0821] Server: Based on the user's viewing history and interests, the server selects the most suitable videos and generates a recommendation list. This list is sent to the user's device.
[0822] Emotion analysis and recommendations using an emotion engine
[0823] 1. The server uses an emotion engine to analyze the user's emotions.
[0824] Server: Analyzes emotions based on the user's facial expressions, tone of voice, viewing history, etc.
[0825] 2. The server reflects the results of the sentiment analysis in its video recommendations.
[0826] Server: Updates the recommendation list taking into account emotional data to provide more of the content the user enjoys.
[0827] Video Ad Optimization
[0828] 1. The server selects video ads based on user characteristics.
[0829] Server: Selects the most suitable advertisement based on the user's age, gender, interests, and emotional analysis results. Also sets the advertisement meta information and playback timing.
[0830] 2. The server delivers ads within the video.
[0831] Server: Evaluates the relevance of the video content and inserts ads at the appropriate time. When a user watches a video, relevant ads are displayed.
[0832] Determining what to monitor
[0833] 1. The server determines whether to monitor the video based on the summary.
[0834] Server: Detects whether the video summary contains inappropriate content (e.g. violence, adult content, etc.), performs a risk assessment, and determines whether manual supervision is required.
[0835] 2. The server will tag the data for manual monitoring as needed.
[0836] Server: Videos that are deemed safe by automatic monitoring are exempt from manual monitoring. If manual monitoring is deemed necessary, the monitoring team is notified.
[0837] Applying an age filter
[0838] 1. The server applies an age filter based on the user's estimated age.
[0839] Server: Estimates the user's age from their registration information and activity history. Detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[0840] 2. The server filters inappropriate content.
[0841] Server: Content that is not appropriate for the specified age group is not played on the user's device, thereby preventing minors from accessing inappropriate content.
[0842] Specific examples
[0843] Example 1: User-Submitted Pet Videos
[0844] 1. The user (device) uploads a video of their dog playing at home.
[0845] 2. The server receives the video, performs quality checks, and generates metadata.
[0846] 3. The server analyzes the video and generates keywords such as "pet," "dog," and "play."
[0847] 4. The server uses these keywords to recommend pet videos relevant to the user.
[0848] 5. The server analyzes the user's emotions, detects whether the user is happy, and recommends more pet videos based on that.
[0849] Example 2: Fashion-related video ads
[0850] 1. The server identifies that the user has fashion-related interests.
[0851] 2. The server uses the emotion engine to confirm that the user has positive emotions toward the fashion content.
[0852] 3. The server selects fashion-related video ads and inserts them into the video at the appropriate time.
[0853] 4. The device displays relevant ads while you watch fashion videos.
[0854] Example 3: Reducing monitoring costs
[0855] 1. The server verifies the summary of the posted video for inappropriate content.
[0856] 2. If the server determines that there is no problem, the video will be automatically made public and will no longer be subject to manual monitoring.
[0857] Note
[0858] This system utilizes an emotion engine to provide content tailored to user characteristics, optimize advertising revenue, and reduce manual monitoring costs. In addition, by taking into account the user's emotional state, an improved user experience is expected.
[0859] The processing flow will be explained below.
[0860] Step 1:
[0861] A user uses a terminal to upload a video file and its meta information (title, description, etc.) to a server.
[0862] Step 2:
[0863] The server checks the format of the received video file and converts it to a standard format if necessary, as well as checking the video quality and generating metadata (file size, duration, etc.).
[0864] Step 3:
[0865] The server begins analyzing the video content using AI technology, specifically by analyzing the frames of the video and identifying key scenes and important objects.
[0866] Step 4:
[0867] The server performs audio analysis to identify speakers and extract key phrases, and generates a summary of the video based on the results of this analysis.
[0868] Step 5:
[0869] Based on the generated summary, the server assigns tags and keywords to the video and classifies it into a specific category, thereby categorizing the video.
[0870] Step 6:
[0871] The server selects appropriate videos based on the user's viewing history and interests, and generates a recommendation list for the user. This list is sent to the device.
[0872] Step 7:
[0873] The server uses an emotion engine to analyze the user's emotions. Specifically, it collects data such as the user's facial expressions, voice, and viewing history to evaluate their emotional state.
[0874] Step 8:
[0875] The server updates the recommendation list based on the results of the sentiment analysis. Specifically, it prioritizes recommendations of content that users enjoy watching.
[0876] Step 9:
[0877] The server selects video ads based on user characteristics and sentiment analysis results, selects the most suitable ad, and also sets its meta information and playback timing.
[0878] Step 10:
[0879] The server evaluates the relevance of the video content and inserts advertisements at the appropriate time, and the device displays the advertisements.
[0880] Step 11:
[0881] The server automatically determines what to monitor based on the video summary, and detects inappropriate content (e.g., violence, adult content, etc.) from the summary.
[0882] Step 12:
[0883] The server performs a risk assessment and, if automated monitoring determines there are no issues, the video is made public. If manual monitoring is deemed necessary, the monitoring team is notified.
[0884] Step 13:
[0885] The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[0886] Step 14:
[0887] The server filters content that is inappropriate for the specified age and restricts playback on the device, preventing minors from accessing inappropriate content.
[0888] Example 2
[0889] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0890] In today's video streaming services, it is extremely difficult to efficiently process the large number of videos posted by users and provide relevant recommended videos to users. Furthermore, the task of appropriately filtering inappropriate content and ensuring user safety is a heavy burden. Furthermore, there is a demand for sophisticated recommendations and ad delivery based on user sentiment, and a system that can realize these elements in an integrated manner is needed.
[0891] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for uploading videos from a terminal, a means for checking the quality of the received videos and generating metadata, a means for analyzing the content of the videos and generating summaries, a means for categorizing the videos based on the summaries generated by the server and assigning tags and keywords, a means for adding the videos to a recommendation list based on user characteristics, a means for improving the video recommendation function using an emotion engine, a means for selecting and delivering optimal advertisements based on user characteristics, a means for determining whether automatic or manual monitoring is required based on the content of the videos, and a means for applying an age filter based on the estimated age of the user. This makes it possible to efficiently process videos posted by users, provide highly relevant recommendations, deliver optimized advertisements, and provide a safer usage environment.
[0892] "Terminal" is a general term for electronic devices that users use to upload video files and their meta information to a server, and includes, for example, smartphones and personal computers.
[0893] "Server" refers to a central computer system that receives videos uploaded by users and performs processes such as analysis, conversion, summary generation, recommendation and advertisement delivery.
[0894] "Uploading a video" refers to the process in which a user sends a video file from their own device to a server.
[0895] "Quality confirmation" is the process in which the server checks the quality of the video file received, including the format, resolution, and playback time, and converts it into a standard format.
[0896] "Metadata generation" refers to the process by which the server generates additional information, such as file size, playback time, and resolution, based on the received video file.
[0897] "Video content analysis" is the process in which the server uses AI technology to analyze video frames and audio to identify key scenes and important objects.
[0898] "Summary generation" refers to the process in which the server extracts key scenes and important information based on the analysis of the video content and generates a compact summary.
[0899] "Categorization" is the process of classifying videos into specific categories based on server-generated summaries and assigning tags and keywords.
[0900] A "recommendation list" is a list of videos that the server determines to be optimal based on the user's characteristics and provides them to the user.
[0901] An "emotion engine" refers to software or a system that has the function of analyzing emotions from a user's facial expressions, tone of voice, viewing history, etc.
[0902] "Advertisement selection and delivery" is the process in which the server selects the most appropriate advertisement based on user characteristics and sentiment analysis results, and delivers it to the user within the video at the appropriate time.
[0903] "Automatic monitoring" is a function in which the server analyzes the video content and automatically detects and evaluates inappropriate content based on predetermined criteria.
[0904] "Human monitoring" is a process in which a team of human monitors conducts additional verification on videos that are deemed high risk by automated monitoring.
[0905] "Age filter" refers to a feature that restricts content that is displayed or played based on the estimated age of the user.
[0906] This invention is a system that efficiently processes videos posted by users and delivers highly relevant recommendations and optimized advertisements. This system is mainly built based on communications between terminals, a server, and users, and specific embodiments are described below.
[0907] Video Ingestion and Analysis
[0908] 1. The device sends the video uploaded by the user to the server.
[0909] Users send video files and their meta information (title, description, etc.) to the server from their devices such as smartphones or PCs.
[0910] 2. The server receives the video, checks the quality, and generates metadata.
[0911] The server checks the format of the video file, converts it to a standard format if necessary, checks the video quality, and generates metadata (file size, duration, etc.).
[0912] 3. The server uses AI technology to analyze the video content and generate a summary.
[0913] The server analyzes the video frames to identify key scenes and important objects, and performs audio analysis to identify speakers and extract important phrases. Based on this information, a video summary is generated.
[0914] Categorizing the video
[0915] 1. The server categorizes the video based on the generated summary.
[0916] The server classifies videos into specific categories based on tags and keywords extracted from the summaries, and also assigns these tags and keywords to the video data.
[0917] 2. The server creates a recommendation list appropriate for each user.
[0918] The server selects the most suitable videos based on the user's viewing history and interests, and generates a recommendation list, which is then sent to the user's device.
[0919] Emotion analysis and recommendations using an emotion engine
[0920] 1. The server analyzes the user's emotions using an emotion engine.
[0921] The server analyzes the user's emotions based on facial expressions, tone of voice, viewing history, etc.
[0922] 2. The server reflects the results of the sentiment analysis in its video recommendations.
[0923] The server takes the emotional data into account and updates the recommendation list to provide more of the content the user enjoys.
[0924] Video Ad Optimization
[0925] 1. The server selects video ads based on user characteristics.
[0926] The server selects the most suitable advertisement based on the user's age, gender, interests, and emotional analysis results, and also sets the advertisement's meta information and playback timing.
[0927] 2. The server delivers ads within the video.
[0928] The server evaluates the relevance of the video content and inserts ads at the appropriate time, so that relevant ads are displayed when the user watches the video.
[0929] Determining what to monitor
[0930] 1. The server determines whether to monitor the video based on the summary.
[0931] The server detects whether the video summary contains inappropriate content (e.g., violence, adult content, etc.), performs a risk assessment, and determines whether manual supervision is necessary.
[0932] 2. The server will tag the data for manual monitoring as needed.
[0933] Videos that are deemed safe by automatic server monitoring will not be subject to manual monitoring. If manual monitoring is deemed necessary, the monitoring team will be notified.
[0934] Applying an age filter
[0935] 1. The server applies an age filter based on the user's estimated age.
[0936] The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[0937] 2. The server filters inappropriate content.
[0938] The server prevents content that is not appropriate for the specified age from being played on the user's device, thereby preventing minors from accessing inappropriate content.
[0939] Specific examples
[0940] Example 1: User-Submitted Pet Videos
[0941] 1. The user (device) uploads a video of their dog playing at home.
[0942] 2. The server receives the video, performs quality checks, and generates metadata.
[0943] 3. The server analyzes the video and generates keywords such as "pet," "dog," and "play."
[0944] 4. The server uses these keywords to recommend relevant pet videos to the user.
[0945] 5. The server performs sentiment analysis on the user, detects whether the user is happy, and recommends more pet videos based on that.
[0946] Example 2: Fashion-related video ads
[0947] 1. The server identifies that the user has fashion-related interests.
[0948] 2. The server uses an emotion engine to confirm that the user has positive emotions toward the fashion content.
[0949] 3. The server selects fashion-related video ads and inserts them into the video at the appropriate time.
[0950] 4. The device displays relevant ads while you watch fashion videos.
[0951] Example 3: Reducing monitoring costs
[0952] 1. The server verifies the summary of the posted video for inappropriate content.
[0953] 2. If the server determines that there is no problem, the video will be automatically made public and will no longer be subject to manual monitoring.
[0954] Prompt Sentence Examples
[0955] 1. Analyze the video content and extract key scenes and important phrases.
[0956] 2. Generate a recommendation list based on the user's viewing history.
[0957] 3. Use a sentiment engine to analyze user sentiment and provide appropriate content.
[0958] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0959] Step 1:
[0960] The device sends the video file and its meta information (title, description, etc.) to the server.
[0961] Input: A user uploads a video file and its meta information from their device.
[0962] Processing: The device sends the video file to the server.
[0963] Output: The video file and associated meta information are saved on the server.
[0964] Step 2:
[0965] The server checks the format of the video file received and converts it to a standard format if necessary.
[0966] Input: The submitted video file.
[0967] Processing: The server analyzes the format of the video file and converts it to a standard format (e.g., adjusting the resolution, converting the file format).
[0968] Output: Video files converted to standard format.
[0969] Step 3:
[0970] The server checks the quality of the received video and generates metadata.
[0971] Input: Video files converted to standard formats.
[0972] Processing: The server checks the quality of the video (e.g., resolution, duration, file size, etc.) and generates metadata (e.g., resolution, duration, file size, frame rate).
[0973] Output: Quality-checked video files and their metadata.
[0974] Step 4:
[0975] The server uses AI technology to analyze the video content and generate a summary.
[0976] Input: Quality-checked video files and metadata.
[0977] Processing: The server uses generative AI models to analyze video frames and audio, extract key scenes and phrases, identify speakers in the video, and generate a summary.
[0978] Output: Summary of video content and analysis results.
[0979] Step 5:
[0980] The server categorizes the video based on the generated summary and assigns tags and keywords.
[0981] Input: Summary of video content and analysis results.
[0982] Processing: Based on the summary, the server assigns tags and keywords (e.g., "pets," "dogs," "play," etc.) to the video and classifies it into a specific category.
[0983] Output: Video data with tags and keywords.
[0984] Step 6:
[0985] The server creates a recommendation list suitable for each user.
[0986] Input: Video data with tags and keywords, user viewing history.
[0987] Processing: The server selects relevant videos based on the user's viewing history and interests and generates a recommendation list.
[0988] Output: A list of recommendations suitable for the user.
[0989] Step 7:
[0990] The server analyzes the user's emotions using an emotion engine.
[0991] Input: User's viewing history, facial expressions, and tone of voice.
[0992] Processing: The server uses an emotion engine to analyze the user's facial expressions and tone of voice and extract emotional data.
[0993] Output: User sentiment analysis results.
[0994] Step 8:
[0995] The server reflects the results of the sentiment analysis in its video recommendations.
[0996] Input: User sentiment analysis results, recommendation list.
[0997] Processing: The server takes into account the emotion data and updates the recommendation list to provide more content that the user enjoys.
[0998] Output: An updated recommendation list reflecting the sentiment data.
[0999] Step 9:
[1000] The server selects video advertisements based on user characteristics.
[1001] Input: User's age, gender, interests, and sentiment analysis results.
[1002] Processing: The server selects the most suitable advertisement based on the user's characteristics and sets the advertisement's meta information and playback timing.
[1003] Output: Advertising data based on user characteristics.
[1004] Step 10:
[1005] The server delivers ads within the video.
[1006] Input: Advertising data based on user characteristics, video being watched.
[1007] Processing: The server evaluates the relevance of the video content and inserts advertisements at the appropriate time.
[1008] Output: Ads inserted during viewing.
[1009] Step 11:
[1010] The server determines whether the video should be monitored based on the summary.
[1011] Input: A summary of the video content.
[1012] Processing: The server detects inappropriate content (e.g., violence, adult content) from the video summary and performs a risk assessment.
[1013] Output: The result of automated or manual monitoring.
[1014] Step 12:
[1015] The server will tag the monitors manually as needed.
[1016] Input: Verdict of automated or manual monitoring.
[1017] Processing: The server automatically monitors and publishes videos that it determines to be problem-free, and notifies the monitoring team if it determines that there is a problem.
[1018] Output: published video or notification to surveillance team.
[1019] Step 13:
[1020] The server applies an age filter based on the user's estimated age.
[1021] Input: User registration information, activity history, video summary.
[1022] Processing: The server estimates the user's age and detects age-restricted keywords (e.g., "violence" and "adult") from the video summary.
[1023] Output: The filtered content.
[1024] Step 14:
[1025] The server filters inappropriate content.
[1026] Input: Content after age filter applied.
[1027] Processing: The server filters content that is inappropriate for the specified age group and prevents it from being played on the user's device.
[1028] Output: Content that is restricted from being accessed by minors.
[1029] (Application example 2)
[1030] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1031] In modern content distribution services, the large number of users posting videos necessitates efficient video summarization and the creation of recommendation lists tailored to each user. It is also important to optimize ad delivery, reduce monitoring costs, and provide content based on user characteristics. However, no system exists that can simultaneously address all of these challenges. In particular, video recommendations and ad delivery that take user emotions into account have yet to be realized, so an effective solution is needed.
[1032] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1033] In this invention, the server includes: a means for uploading videos from a terminal; a means for checking the quality of the received videos and generating metadata; a means for analyzing the video content and generating a summary; a means for categorizing the videos based on the generated summary and assigning tags and keywords; a means for adding the videos to a recommendation list based on user characteristics; a means for selecting and assigning optimal advertisements based on user characteristics; a means for determining whether automatic or manual monitoring is required based on the video content; a means for applying an age filter based on the estimated age of the user; a means for analyzing the user's emotions using emotion recognition technology and updating the video recommendation list based on the analysis; and a means for inserting video advertisements related to specific video categories into videos at appropriate times. This enables efficient creation of video summaries and recommendation lists, optimization of advertisement distribution, reduction of monitoring costs, and provision of advanced content that takes user emotions into consideration.
[1034] A "terminal" is an electronic device such as a smartphone or computer that a user uses to upload videos.
[1035] "Server" refers to a computer system on a network that receives videos, analyzes them, generates summaries, categorizes them, creates recommendation lists, selects advertisements, determines whether they are monitored automatically or manually, applies age filters, analyzes emotions, and delivers advertisements.
[1036] "Quality check" is a process in which the server checks the format, image quality, sound quality, etc. of the uploaded video and converts it to a standard format if necessary.
[1037] "Metadata generation" is the process by which the server generates the video's file size, playback time, and other related information.
[1038] "Summary generation" is a process in which the server analyzes the content of a video and creates a short summary of the video based on key scenes, important objects, and audio.
[1039] "Categorization" is a process in which the server classifies videos into specific categories based on the summary information generated and assigns related tags and keywords.
[1040] A "recommendation list" is a list in which the server selects specific videos based on the user's viewing history and interests and provides them to the user.
[1041] "Advertisement selection" is a process in which the server selects the most suitable advertisement based on the user's age, gender, interests, and emotional analysis results, and sets its meta information and playback timing.
[1042] "Automatic monitoring" is a process in which the server detects whether a video contains inappropriate content based on its summary information and determines whether manual monitoring is necessary.
[1043] "Age filter" is a process in which the server controls the viewing of inappropriate content based on the estimated age of the user.
[1044] "Emotion recognition technology" is a technology that allows the server to analyze the user's facial expressions, tone of voice, etc. to understand the user's emotional state.
[1045] A "video category" is a specific genre or theme that the server uses to categorize videos, such as pets, fashion, cooking, etc.
[1046] The system for implementing this invention is configured by linking user terminals, a server, and a network, and enables efficient video summarization, improved recommendation functionality, optimized ad distribution, reduced monitoring costs, and a safe usage environment.
[1047] First, a user uploads a video from a device such as a smartphone or PC. This device is used to send the video taken by the user to the server.
[1048] The server temporarily stores the received video and checks its quality. Specifically, it checks the video format and converts it to a standard format if necessary. It also checks the video's image quality and sound quality and generates metadata (such as file size and playback time). This process uses video editing libraries (e.g., OpenCV) and audio analysis libraries.
[1049] The server then uses AI technology to analyze the video content. This analysis involves identifying important scenes in each frame of the video and extracting key phrases from the audio data. The AI technology used here utilizes computer vision for video analysis and a speech recognition engine (e.g., Google Speech-to-Text API) for audio analysis. The generated summary is then compiled in a format that is easy for the user to view.
[1050] The server then categorizes the video based on the generated summary and assigns relevant tags and keywords to it, using natural language processing (NLP) techniques based on information extracted from the summary.
[1051] The server then generates a recommendation list based on the user's characteristics. The most suitable videos are selected based on the user's viewing history and interests, and this list is sent to the user's device. The recommendation engine uses machine learning models to analyze user behavior data.
[1052] Furthermore, the server uses emotion recognition technology to analyze the user's emotions. This analysis involves analyzing the user's facial expressions and tone of voice. The emotion engine uses, for example, the FER (Facial Emotion Recognition) library. Based on the analysis results, the video recommendation list according to the emotion is updated.
[1053] The server selects the most suitable advertisement based on the user's characteristics and inserts it into the video at the appropriate time. Ad selection takes into account age, gender, interests, and emotional analysis results. Ad delivery also uses an algorithm that evaluates the relevance of the advertisement to the video content.
[1054] Finally, the server has a means to determine whether to monitor automatically or manually based on the video content. The automatic monitoring system analyzes the summary of the uploaded video to detect inappropriate content, such as illegal content. If necessary, it notifies the manual monitoring team.
[1055] A concrete example is the automatic recommendation of pet videos. When a user uploads a video of their dog playing at home, the system receives and analyzes the video, generating tags such as "pet," "dog," and "play." Based on the user's past viewing history and sentiment analysis data, the system recommends new pet videos.
[1056] Example of an input prompt for a generative AI model:
[1057] "Enter the video title, description, and uploaded file path. Example: Title: 'Dogs playing at home', Description: 'Dogs playing filmed at home', File path: 'path / to / video.mp4'"
[1058] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1059] Step 1:
[1060] Videos are uploaded from the device
[1061] A user shoots a video using a smartphone or computer and accesses the upload page. The prompt message is "Please enter the video title, description, and uploaded file path. Example: Title: 'Dog playing at home', Description: 'Dog playing filmed at home', File path: 'path / to / video.mp4'." When the user enters the title, description, and video file and presses the upload button, the video file and metadata are sent to the server. The input is the video file and its metadata specified by the user, and the output is the video file and its metadata stored on the server.
[1062] Step 2:
[1063] The server checks the quality of the received video and generates metadata.
[1064] The server temporarily stores the received video file and checks its format, image quality, and sound quality. For example, it uses OpenCV to check the video resolution and pixel quality and converts it to a standard format if necessary. It also uses an audio analysis library to check the sound quality. Metadata such as file size, playback time, and resolution are generated. The input is the received video file, and the output is the generated metadata.
[1065] Step 3:
[1066] The server analyzes the video content and generates a summary
[1067] The server uses AI technology to analyze the video content. Each frame of the video is analyzed using computer vision technology to identify key scenes. For example, OpenCV and deep learning models are used to detect key objects and movements. The audio is converted to text using the Google Speech-to-Text API, and key phrases are extracted. A summary is generated based on this information. The input is each frame of the video and audio data, and the output is a video summary.
[1068] Step 4:
[1069] The server categorizes the video based on the generated summary and assigns tags and keywords.
[1070] The server classifies videos into specific categories based on information obtained from the video summaries and assigns associated tags and keywords. For example, natural language processing techniques are used to extract keywords obtained from the summaries and map them to predefined categories. The input is the generated video summaries, and the output is the classified video categories and assigned tags and keywords.
[1071] Step 5:
[1072] The server adds videos to the recommendation list based on the user's characteristics.
[1073] The server analyzes the user's viewing history, interests, and past viewing data to select the most suitable videos and generate a recommendation list. Here, a machine learning model analyzes the user's behavioral data and lists highly relevant videos. For example, a collaborative filtering algorithm is used. The input is the user's viewing history and interest data, and the output is a recommendation list.
[1074] Step 6:
[1075] The server uses emotion recognition technology to analyze the user's emotions and updates the video recommendation list based on that.
[1076] The server performs facial expression analysis and voice tone analysis to recognize the user's emotions. For example, it uses the FER (Facial Emotion Recognition) library to analyze facial expressions in frames and identify the user's emotions. Based on the results, it adds more types of videos that the user enjoys to the recommendation list. The input is each frame of the video and audio data, and the output is an updated recommendation list.
[1077] Step 7:
[1078] The server selects the most suitable advertisement based on the user's characteristics and inserts it into the video at the appropriate time.
[1079] The server selects the most suitable advertisement based on the user's age, gender, interests, and emotional analysis results, and determines the timing of its insertion into the video. For example, collaborative filtering and emotional analysis data can be combined to identify the user's areas of interest, select highly relevant advertisements, and incorporate them into the video frame. The input is the user's characteristics and the results of the emotional analysis, and the output is a video containing the inserted advertisement.
[1080] Step 8:
[1081] The server determines whether to monitor automatically or manually based on the video content.
[1082] The server uses the generated summary to check for inappropriate content in the video using an automated monitoring system. For example, it uses content analysis algorithms to detect problematic scenes or audio, performs a risk assessment, and notifies a human monitoring team if necessary. The input is the generated video summary, and the output is the monitoring decision.
[1083] Step 9:
[1084] The server applies an age filter based on the user's estimated age.
[1085] The server estimates the user's age based on their registration information and activity history, and restricts access to inappropriate content. For example, it extracts age-related keywords from the summary information and filters videos that do not match the user's age. The input is the user's estimated age and the video summary, and the output is a filtered list of videos.
[1086] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1087] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1088] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1089] [Third embodiment]
[1090] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1091] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1092] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1093] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1094] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1095] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1096] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1097] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1098] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1099] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1100] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1101] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1102] This invention is a system that efficiently summarizes videos posted by users and, based on the summaries, enables recommendation functions, advertisement distribution, reduced monitoring costs, and a safe usage environment. This system is built on communications between terminals, servers, and users. Below, the processing of the system's program is explained in natural language, with specific examples.
[1103] Video Ingestion and Analysis
[1104] 1. The device sends the video uploaded by the user to the server.
[1105] Device: The user uploads the video file and its meta information (title, description, etc.) to the server from a device such as a smartphone or PC.
[1106] 2. The server receives the video and performs the ingest process.
[1107] Server: Checks the format of the video file and converts it to a standard format if necessary. It also checks the video quality and generates metadata (file size, duration, etc.).
[1108] 3. The server uses AI technology to analyze the video content and generate a summary.
[1109] Server: Analyzes video frames to identify key scenes and important objects. Also, analyzes audio to identify speakers and extract important phrases. Then, it generates a video summary based on this information.
[1110] Categorizing the video
[1111] 1. The server categorizes the video based on the generated summary.
[1112] Server: Classifies videos into specific categories based on tags and keywords extracted from the summaries, and assigns these tags and keywords to the video data.
[1113] 2. The server creates a recommendation list suitable for each user.
[1114] Server: Based on the user's viewing history and interests, the server selects the most suitable videos and generates a recommendation list. This list is sent to the user's device.
[1115] Video Ad Optimization
[1116] 1. The server selects video ads based on user characteristics.
[1117] Server: Selects the most suitable advertisement based on the user's characteristics such as age, gender, and interests. Also sets the advertisement meta information and playback timing.
[1118] 2. The server delivers ads within the video.
[1119] Server: Evaluates the relevance of the video content and inserts ads at the appropriate time. When a user watches a video, relevant ads are displayed.
[1120] Determining what to monitor
[1121] 1. The server determines whether to monitor the video based on the summary.
[1122] Server: Detects whether the video summary contains inappropriate content (e.g. violence, adult content, etc.), performs a risk assessment, and determines whether manual supervision is required.
[1123] 2. The server will tag the data for manual monitoring as needed.
[1124] Server: Videos that are deemed safe by automatic monitoring are exempt from manual monitoring. If manual monitoring is deemed necessary, the monitoring team is notified.
[1125] Applying an age filter
[1126] 1. The server applies an age filter based on the user's estimated age.
[1127] Server: Estimates the user's age from their registration information and activity history. Detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[1128] 2. The server filters inappropriate content.
[1129] Server: Content that is not appropriate for the specified age group will not be played on the user's device, thereby preventing minors from accessing inappropriate content.
[1130] Specific examples
[1131] Example 1: User-Submitted Pet Videos
[1132] 1. The user (device) uploads a video of their dog playing at home.
[1133] 2. The server receives the video, performs quality checks, and generates metadata.
[1134] 3. The server analyzes the video and generates keywords such as "pet," "dog," and "play."
[1135] 4. The server uses these keywords to recommend pet videos relevant to the user.
[1136] Example 2: Fashion-related video ads
[1137] 1. The server identifies that the user has fashion-related interests.
[1138] 2. The server selects fashion-related video ads and inserts them into the video at the appropriate time.
[1139] 3. The device displays relevant ads while you watch fashion videos.
[1140] Example 3: Reducing monitoring costs
[1141] 1. The server verifies the summary of the posted video for inappropriate content.
[1142] 2. If the server determines that there is no problem, the video will be automatically made public and will no longer be subject to manual monitoring.
[1143] This system enables the provision of content tailored to user characteristics, optimization of advertising revenue, and reduction of human monitoring costs.
[1144] The processing flow will be explained below.
[1145] Step 1:
[1146] A user uses a terminal to upload a video file and its meta information (title, description, etc.) to a server.
[1147] Step 2:
[1148] The server checks the format of the video file received and converts it to a standard format if necessary, checks the video quality, and generates metadata (file size, playback time, etc.).
[1149] Step 3:
[1150] The server begins analyzing the video content using AI technology, specifically by analyzing the frames of the video and identifying key scenes and important objects.
[1151] Step 4:
[1152] The server performs audio analysis to identify speakers and extract key phrases, and generates a video summary based on the results of this analysis.
[1153] Step 5:
[1154] Based on the generated summary, the server assigns tags and keywords to the video and classifies it into a specific category, thereby categorizing the video.
[1155] Step 6:
[1156] The server selects appropriate videos based on the user's viewing history and interests, and generates a recommendation list for the user. This list is sent to the device.
[1157] Step 7:
[1158] The server selects the most suitable video ad based on user characteristics (age, gender, interests, etc.) and also sets the ad's meta information and playback timing.
[1159] Step 8:
[1160] The server evaluates the relevance of the video content and inserts ads at the appropriate time, and displays relevant ads when the device is watching the video.
[1161] Step 9:
[1162] Based on the summary, the server determines whether the video is suitable for automatic monitoring or requires manual monitoring. Inappropriate content (e.g., violence, adult content, etc.) is detected.
[1163] Step 10:
[1164] The server performs a risk assessment and, if it determines that there are no problems with automated monitoring, it makes the video public. If it determines that manual monitoring is necessary, it notifies the monitoring team.
[1165] Step 11:
[1166] The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the summary.
[1167] Step 12:
[1168] The server filters content that is inappropriate for the specified age and prevents the video from being played on the device, thereby preventing minors from accessing inappropriate content.
[1169] Example 1
[1170] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1171] Modern video streaming services require efficient analysis and classification of the vast amount of videos uploaded by users, and the provision of appropriate recommendations and advertisements. However, manual video analysis and recommendation generation requires a great deal of effort and time, resulting in high operational costs, making efficient and automated processing methods necessary. Automation is also required for providing personalized advertisements based on user characteristics, filtering content suitable for underage users, and monitoring video content. An efficient and reliable system is needed to solve these challenges.
[1172] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1173] In this invention, the server includes a means for uploading videos from a terminal, a means for converting the format of the received video, checking the quality, and generating metadata, a means for analyzing the content of the video using a generation AI model and generating a summary, a means for categorizing the video based on the generated summary and assigning tags and keywords, a means for adding the video to a recommendation list based on user characteristics, a means for selecting and delivering optimal advertisements based on user characteristics, a means for determining whether automatic or manual monitoring is required based on the video content, and a means for applying an age filter based on the estimated age of the user. This enables efficient analysis of videos uploaded by users, provision of appropriate recommendations and advertisements, automated content monitoring, and age-appropriate filtering.
[1174] A "terminal" is a device used by a user, such as a smartphone or computer, used to upload videos.
[1175] The "server" is a central device that receives, analyzes, and processes video data over the network, and plays a central role in this system.
[1176] A "video" is a file containing a series of video and audio files uploaded by a user.
[1177] "Format conversion" is the process of converting received video files into a standard format.
[1178] "Quality check" is the process of checking the quality of the received video, such as its resolution and playback time.
[1179] "Metadata" refers to various information related to a video file (e.g., file size, playback time, resolution, etc.).
[1180] A "generative AI model" is an artificial intelligence technology used to analyze video content and generate summaries.
[1181] "Video analysis" is the process of analyzing video frames and audio to identify important scenes and objects.
[1182] "Summary" refers to a short summary of the main content extracted from the results of video analysis.
[1183] "Tags" refer to keywords or labels that are added to indicate the content of a video.
[1184] "Keywords" refer to words or short phrases related to the video content.
[1185] A "recommendation list" is a list of videos selected based on the user's characteristics and interests.
[1186] "Advertisement" refers to information including product promotions and service information, which is inserted into videos so that users can view them.
[1187] "Automatic monitoring" is the process of automatically checking video content to detect inappropriate content.
[1188] "Manual monitoring" is a manual monitoring process carried out by specialized staff when automatic monitoring is difficult to judge.
[1189] "Estimated age" is the age calculated based on the user's registration information and activity history.
[1190] An "age filter" is a mechanism that provides only appropriate content based on the user's estimated age.
[1191] This invention is a system that efficiently summarizes videos posted by users and provides recommendation functions, advertisement distribution, reduced monitoring costs, and a safe usage environment based on the summaries. This system is built on communications between terminals, a server, and users.
[1192] Video Ingestion and Analysis
[1193] When a user uploads a video to a server from a device, the user uses a device such as a smartphone or PC. Specifically, the user sends the video file and its meta information (title, description, etc.) to the server through an interface provided by the device.
[1194] The server receives video files uploaded by users. The received video files are first converted into a different format. Software such as ffmpeg is used for this process. Next, the quality of the video is checked and metadata (file size, resolution, playback time, etc.) is generated based on that quality.
[1195] The server uses a generative AI model to analyze the video content and generate a summary. Specifically, it uses an object detection model such as YOLO (You Only Look Once) to identify key scenes and important objects, and then performs audio analysis using the Google Cloud Speech-to-Text API. Then, based on this information, it uses a natural language generation model (e.g., GPT-3) to generate a video summary.
[1196] Video categorization and recommendations
[1197] The server categorizes the videos based on the generated summaries. The videos are classified into specific categories based on tags and keywords extracted from the summaries. The classified videos are then assigned tags and keywords to improve searchability and filtering capabilities.
[1198] The server then generates a recommendation list based on the user's characteristics. Taking into account the user's viewing history and interests, the server selects the most suitable videos to create the recommendation list. This recommendation list is then sent to the user's device, where the user can watch the recommended videos through an interface.
[1199] Ad optimization and delivery
[1200] The server selects the most suitable advertisement based on the video summary and user characteristics. Specifically, it selects advertisements based on the user's characteristics such as age, gender, and interests, and sets the meta information and playback timing. The server references a database provided by the advertiser to select advertisements.
[1201] The server delivers advertisements at appropriate times within the video. For example, while watching a video about pets, it displays an advertisement for related pet food. This can be expected to have a high advertising effect.
[1202] Determining who should be monitored and applying age filters
[1203] The server determines whether a video should be subject to automated or manual monitoring based on the summary. The video summary is used to detect whether it contains inappropriate content and perform a risk assessment. Based on this, the process of determining whether to subject the video to automated or manual monitoring uses natural language processing technology.
[1204] Furthermore, the server applies an age filter based on the user's estimated age. The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the video summary. Inappropriate content is filtered so that it cannot be played on devices belonging to underage users. Based on these filtering rules, only appropriate videos are provided.
[1205] Specific examples
[1206] Example 1: User-Submitted Pet Videos
[1207] A user uploads a video of their dog playing at home. The server receives the video, checks the quality, and generates metadata. The server then analyzes the video and generates keywords such as "pet," "dog," and "play." The server then recommends relevant pet videos to the user based on these keywords.
[1208] Example 2: Fashion-related video ads
[1209] The server identifies the user's interest in fashion. The server then selects fashion-related video ads and inserts them into the video at the appropriate time. The relevant ads are then displayed on the user's device while they are watching the fashion video.
[1210] Example 3: Reducing monitoring costs
[1211] The server verifies the summary of the posted video to see if it contains inappropriate content. If it determines there are no problems, it automatically publishes the video and removes it from the scope of human monitoring.
[1212] Prompt Sentence Examples
[1213] "Recommend other videos related to this pet video."
[1214] "Insert fashion-related ads at the best possible time."
[1215] Please determine if this video is inappropriate.
[1216] This system makes it possible to provide content tailored to user characteristics, optimize advertising revenue, and reduce manual monitoring costs.
[1217] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1218] Step 1:
[1219] The device sends the video uploaded by the user to the server. The user clicks the upload button through the device interface, selects the video file (e.g., mp4, avi) and its meta information (title, description), and sends it to the server.
[1220] Input: Video file and meta information
[1221] Output: Video file and meta information sent to the server
[1222] Step 2:
[1223] The server receives the video, converts the format, checks the quality, and generates metadata. The server checks the format of the received video file and converts it to a standard format (e.g., mp4). This process uses tools such as ffmpeg. It also checks the quality of the video, such as resolution, playback time, and file size, and generates this information as metadata.
[1224] Input: Video file and meta information sent to the server
[1225] Output: Format converted video file and generated metadata
[1226] Step 3:
[1227] The server uses a generative AI model to analyze the video content and generate a summary. Specifically, it uses an object detection model such as YOLO to analyze video frames and identify key scenes and important objects. It also performs audio analysis using the Google Cloud Speech-to-Text API to identify speakers and extract key phrases. It then uses this data to generate a video summary using a natural language generation model (e.g., GPT-3).
[1228] Input: Format converted video file and metadata
[1229] Output: Video summary
[1230] Step 4:
[1231] The server categorizes the video based on the generated summary and assigns tags and keywords. Based on the information obtained from the summary, the server classifies the video into an appropriate category (e.g., pets, sports, entertainment), and assigns tags and keywords related to that category to the video data.
[1232] Input: Video summary
[1233] Output: Categorized videos with tags and keywords
[1234] Step 5:
[1235] The server adds videos to a recommendation list based on user characteristics. The server analyzes the user's viewing history and interests, selects the most suitable videos, and generates a recommendation list. This list is sent to the user's device.
[1236] Input: User's viewing history and interests, as well as categorized videos, tags, and keywords
[1237] Output: Recommendation list based on user characteristics
[1238] Step 6:
[1239] The server selects and delivers the most suitable advertisement based on the user's characteristics. The server selects the appropriate advertisement based on the user's characteristics information such as age, gender, and interests. Specifically, it retrieves the advertisement that best suits the user's characteristics from the advertisement database and sets its meta information and playback timing.
[1240] Input: User characteristics (age, gender, interests, etc.) and advertising database
[1241] Output: Optimal advertisement based on user characteristics
[1242] Step 7:
[1243] The server determines whether the video should be monitored automatically or manually based on its content. The server then checks the video summary for inappropriate content and assesses the content risk. This process uses natural language processing technology.
[1244] Input: Video summary
[1245] Output: Judgment results of automatic or manual monitoring
[1246] Step 8:
[1247] The server applies an age filter based on the user's estimated age. The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords from the video summary. Inappropriate content is filtered so that it cannot be played on devices belonging to minors.
[1248] Input: User registration information, activity history, video summary
[1249] Output: Age-filtered video list
[1250] (Application example 1)
[1251] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1252] In recent years, with the proliferation of video content, a huge amount of videos has been uploaded to online platforms. However, there is a lack of systems that can efficiently summarize videos and recommend appropriate content to users. Furthermore, manually monitoring video content and filtering inappropriate content is extremely time-consuming and costly. Furthermore, optimization of advertisements based on user characteristics is insufficient, and efficient advertisement delivery is required. To address these issues, a system is needed that automates video analysis and recommendation, advertisement optimization, and content monitoring, while providing a safe viewing environment.
[1253] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1254] In this invention, the server includes a means for uploading videos from a terminal, a means for checking the quality of the received videos and generating metadata, a means for analyzing the video content and generating a summary, a means for categorizing the videos based on the generated summary and assigning tags and keywords, a means for adding the videos to a recommendation list based on user characteristics, a means for selecting and assigning optimal advertisements based on user characteristics, a means for determining whether automatic or manual monitoring is required based on the video content, a means for applying an age filter based on the estimated age of the user, a means for saving extracted frame images and transcribing and analyzing audio, a means for providing a function for recommending similar content based on responses from the server, and a means for inserting advertisements and optimizing based on user characteristics as part of a content distribution service. This enables efficient video summarization and recommendation functions, optimal advertisement distribution, reduced monitoring costs, and a safe usage environment.
[1255] - "Device" refers to the device to which the video is uploaded, such as a smartphone or computer.
[1256] "Server" refers to the computing system that processes the received video, checks the quality, generates metadata, and analyzes the video content.
[1257] "Video content analysis" refers to the process of analyzing video frames and audio data to extract summaries, important scenes, and objects.
[1258] "Generating a summary" refers to extracting the main points from the analyzed video data and summarizing them concisely.
[1259] "Categorizing videos" refers to classifying videos into specific categories based on the generated summary and assigning relevant tags and keywords.
[1260] "Add to recommendation list" refers to selecting suitable videos based on the user's viewing history and interests and adding them to the display list.
[1261] "Selecting and delivering ads" refers to selecting the most suitable ads based on user characteristics and inserting them while the video is playing.
[1262] "Determine whether to use automated or manual monitoring" refers to determining whether a video contains inappropriate content based on a summary of the video and deciding whether to use automated or, if necessary, manual monitoring.
[1263] "Applying age filters" refers to restricting the content available to a user based on their estimated age.
[1264] "Save frame images" means saving still images extracted from a video, usually at 1-second intervals.
[1265] "Audio transcription and analysis" refers to the process of converting video audio into text data and extracting key phrases and keywords.
[1266] Providing a "similar content recommendation function" refers to a function that automatically suggests relevant videos based on a user's viewing history and characteristics.
[1267] "Ad insertion and optimization based on user characteristics" refers to inserting ads into videos at the appropriate time and providing the optimal advertising plan based on the user's characteristics.
[1268] This invention provides a system that efficiently summarizes videos posted by users and, based on the summaries, enables recommendation functions, advertisement distribution, reduced monitoring costs, and a safe usage environment. Detailed modes for carrying out the invention are described below.
[1269] Uploading videos from your device
[1270] Users upload video files to the server using devices such as smartphones or PCs. At this time, the video file and its meta information (title, description, etc.) are also sent.
[1271] Server receives and processes video
[1272] The server checks the quality of the received video files, converts them to a standard format if necessary, and generates metadata (file size, playback time, etc.) for the video.
[1273] Video content analysis and summary generation
[1274] The server analyzes the video frames to identify key scenes and important objects, and performs audio analysis to identify speakers and extract key phrases. Based on this information, a video summary is generated. This process includes saving the captured frame images and transcribing and analyzing the audio extracted from the video.
[1275] Video categorization and recommendations
[1276] The server categorizes the videos based on the generated summaries and assigns tags and keywords. Based on the user's characteristics, it selects the most suitable videos and generates a recommendation list. Based on the response from the server, a recommendation function for similar content is provided.
[1277] Ad optimization and insertion
[1278] The server selects the most suitable advertisement based on the user's characteristics and inserts it at a timing relevant to the video content. As part of the content distribution service, it performs advertisement insertion and optimization based on user characteristics.
[1279] Video monitoring and age filtering
[1280] The server determines whether to monitor the content based on the video summary, automatically or manually, and applies an age filter based on the user's estimated age to filter out inappropriate content.
[1281] Hardware and software used
[1282] The following hardware and software are used to implement this system.
[1283] Hardware: Smartphones, PCs
[1284] Software: MoviePy (video editing library), Vosk (voice recognition library), OpenCV (image processing library)
[1285] Specific examples
[1286] Example 1: User-Submitted Pet Videos
[1287] When a user uploads a video of their dog playing at home, the server receives the video, checks the quality, and generates metadata. It then analyzes the video to generate keywords such as "pet," "dog," and "play," and recommends related pet videos.
[1288] Example 2: Fashion-related video ads
[1289] The server identifies that the user is interested in fashion, selects fashion-related video advertisements, and inserts them into the videos at the appropriate time. The relevant advertisements are displayed while the user is watching the fashion video.
[1290] Prompt Sentence Examples
[1291] text
[1292] Analyzes the user-uploaded video "example_video.mp4" and generates a summary. Then, recommends similar videos and inserts relevant ads. Finally, applies supervision and age filters.
[1293] In this way, the present invention makes it possible to provide efficient video summarization and recommendation functions, optimal advertisement distribution, reduced monitoring costs, and a safe usage environment.
[1294] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1295] Step 1:
[1296] A user uses a terminal to upload a video file to a server.
[1297] Input: Video file, meta information (title, description, etc.)
[1298] Output: Video files and meta information are saved on the server
[1299] Specific operation: A user selects a video file and presses the upload button on their smartphone or PC. The device then sends the selected video file and its meta information to the server as an HTTP request.
[1300] Step 2:
[1301] The server checks the quality of the received video file and generates metadata.
[1302] Input: Uploaded video file, meta information
[1303] Output: Video quality evaluation results, metadata (file size, playing time, etc.)
[1304] What it does: The server checks the format of the video file and converts it to a standard format if necessary. It also evaluates the video quality and generates metadata such as file size, duration, and resolution.
[1305] Step 3:
[1306] The server performs frame analysis and audio analysis of the video content and generates a summary.
[1307] Input: Video file
[1308] Output: Video summary, list of important scenes and objects, audio transcription
[1309] Specific operations: Extract key scenes from each frame of a video using MoviePy and OpenCV. Transcribe audio data using Vosk and extract key phrases and keywords. Combine these to generate a video summary.
[1310] Step 4:
[1311] The server categorizes the video based on the generated summary and assigns tags and keywords.
[1312] Input: Video summary
[1313] Output: Video data with tags and keywords
[1314] How it works: The server classifies videos into specific categories based on phrases and keywords extracted from the summaries, automatically assigning tags such as "pets," "dogs," and "play."
[1315] Step 5:
[1316] The server generates a recommendation list based on user characteristics.
[1317] Input: User viewing history, interests, summary data
[1318] Output: Recommendation list for each user
[1319] How it works: The server analyzes the user's past viewing history and interests (e.g., pet videos) and selects relevant videos from the summary data, thereby generating a customized recommendation list for each user.
[1320] Step 6:
[1321] The server selects the most suitable advertisement based on user characteristics and inserts it at the appropriate time within the video.
[1322] Input: User characteristics, video summary, relevant advertising data
[1323] Output: Video with ads
[1324] Specific operation: The server selects the most suitable advertisement based on the user's characteristics such as age, gender, and interests, and inserts the selected advertisement at the relevant timing in the video (e.g., scene change).
[1325] Step 7:
[1326] The server determines which videos to monitor based on the video summary.
[1327] Input: Video summary
[1328] Output: Judgment result of whether the target is automatic or manual monitoring
[1329] How it works: The server evaluates the summary and determines whether it contains inappropriate content. If the automated monitoring finds no problems, the video is left public. If manual monitoring is deemed necessary, the monitoring team is notified.
[1330] Step 8:
[1331] The server applies an age filter based on the user's estimated age.
[1332] Input: User registration information, activity history, video summary
[1333] Output: Filtered video content
[1334] Specific operation: The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords from the video summary. It then filters the video to prevent underage users from accessing inappropriate content.
[1335] Examples of prompt statements
[1336] Analyzes the user-uploaded video "example_video.mp4" and generates a summary. Then, recommends similar videos and inserts relevant ads. Finally, applies supervision and age filters.
[1337] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1338] This invention is a system that efficiently summarizes videos posted by users, improving recommendation functionality, optimizing ad delivery, reducing monitoring costs, and providing a safe usage environment. Furthermore, by combining it with an emotion engine that analyzes user emotions, more advanced recommendations and ad delivery become possible. This system is built on communications between terminals, servers, and users. Below, we explain the system's program processing in natural language, including specific examples.
[1339] Video Ingestion and Analysis
[1340] 1. The device sends the video uploaded by the user to the server.
[1341] Device: The user uploads the video file and its meta information (title, description, etc.) to the server from a device such as a smartphone or PC.
[1342] 2. The server receives the video and performs the ingest process.
[1343] Server: Checks the format of the video file and converts it to a standard format if necessary. It also checks the video quality and generates metadata (file size, duration, etc.).
[1344] 3. The server uses AI technology to analyze the video content and generate a summary.
[1345] Server: Analyzes video frames to identify key scenes and important objects. Also, analyzes audio to identify speakers and extract important phrases. Then, it generates a video summary based on this information.
[1346] Categorizing the video
[1347] 1. The server categorizes the video based on the generated summary.
[1348] Server: Classifies videos into specific categories based on tags and keywords extracted from the summaries, and assigns these tags and keywords to the video data.
[1349] 2. The server creates a recommendation list suitable for each user.
[1350] Server: Based on the user's viewing history and interests, the server selects the most suitable videos and generates a recommendation list. This list is sent to the user's device.
[1351] Emotion analysis and recommendations using an emotion engine
[1352] 1. The server uses an emotion engine to analyze the user's emotions.
[1353] Server: Analyzes emotions based on the user's facial expressions, tone of voice, viewing history, etc.
[1354] 2. The server reflects the results of the sentiment analysis in its video recommendations.
[1355] Server: Updates the recommendation list taking into account emotional data to provide more of the content the user enjoys.
[1356] Video Ad Optimization
[1357] 1. The server selects video ads based on user characteristics.
[1358] Server: Selects the most suitable advertisement based on the user's age, gender, interests, and emotional analysis results. Also sets the advertisement meta information and playback timing.
[1359] 2. The server delivers ads within the video.
[1360] Server: Evaluates the relevance of the video content and inserts ads at the appropriate time. When a user watches a video, relevant ads are displayed.
[1361] Determining what to monitor
[1362] 1. The server determines whether to monitor the video based on the summary.
[1363] Server: Detects whether the video summary contains inappropriate content (e.g. violence, adult content, etc.), performs a risk assessment, and determines whether manual supervision is required.
[1364] 2. The server will tag the data for manual monitoring as needed.
[1365] Server: Videos that are deemed safe by automatic monitoring are exempt from manual monitoring. If manual monitoring is deemed necessary, the monitoring team is notified.
[1366] Applying an age filter
[1367] 1. The server applies an age filter based on the user's estimated age.
[1368] Server: Estimates the user's age from their registration information and activity history. Detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[1369] 2. The server filters inappropriate content.
[1370] Server: Content that is not appropriate for the specified age group is not played on the user's device, thereby preventing minors from accessing inappropriate content.
[1371] Specific examples
[1372] Example 1: User-Submitted Pet Videos
[1373] 1. The user (device) uploads a video of their dog playing at home.
[1374] 2. The server receives the video, performs quality checks, and generates metadata.
[1375] 3. The server analyzes the video and generates keywords such as "pet," "dog," and "play."
[1376] 4. The server uses these keywords to recommend pet videos relevant to the user.
[1377] 5. The server analyzes the user's emotions, detects whether the user is happy, and recommends more pet videos based on that.
[1378] Example 2: Fashion-related video ads
[1379] 1. The server identifies that the user has fashion-related interests.
[1380] 2. The server uses the emotion engine to confirm that the user has positive emotions toward the fashion content.
[1381] 3. The server selects fashion-related video ads and inserts them into the video at the appropriate time.
[1382] 4. The device displays relevant ads while you watch fashion videos.
[1383] Example 3: Reducing monitoring costs
[1384] 1. The server verifies the summary of the posted video for inappropriate content.
[1385] 2. If the server determines that there is no problem, the video will be automatically made public and will no longer be subject to manual monitoring.
[1386] Note
[1387] This system utilizes an emotion engine to provide content tailored to user characteristics, optimize advertising revenue, and reduce manual monitoring costs. In addition, by taking into account the user's emotional state, an improved user experience is expected.
[1388] The processing flow will be explained below.
[1389] Step 1:
[1390] A user uses a terminal to upload a video file and its meta information (title, description, etc.) to a server.
[1391] Step 2:
[1392] The server checks the format of the received video file and converts it to a standard format if necessary, as well as checking the video quality and generating metadata (file size, duration, etc.).
[1393] Step 3:
[1394] The server begins analyzing the video content using AI technology, specifically by analyzing the frames of the video and identifying key scenes and important objects.
[1395] Step 4:
[1396] The server performs audio analysis to identify speakers and extract key phrases, and generates a summary of the video based on the results of this analysis.
[1397] Step 5:
[1398] Based on the generated summary, the server assigns tags and keywords to the video and classifies it into a specific category, thereby categorizing the video.
[1399] Step 6:
[1400] The server selects appropriate videos based on the user's viewing history and interests, and generates a recommendation list for the user. This list is sent to the device.
[1401] Step 7:
[1402] The server uses an emotion engine to analyze the user's emotions. Specifically, it collects data such as the user's facial expressions, voice, and viewing history to evaluate their emotional state.
[1403] Step 8:
[1404] The server updates the recommendation list based on the results of the sentiment analysis. Specifically, it prioritizes recommendations of content that users enjoy watching.
[1405] Step 9:
[1406] The server selects video ads based on user characteristics and sentiment analysis results, selects the most suitable ad, and also sets its meta information and playback timing.
[1407] Step 10:
[1408] The server evaluates the relevance of the video content and inserts advertisements at the appropriate time, and the device displays the advertisements.
[1409] Step 11:
[1410] The server automatically determines what to monitor based on the video summary, and detects inappropriate content (e.g., violence, adult content, etc.) from the summary.
[1411] Step 12:
[1412] The server performs a risk assessment and, if automated monitoring determines there are no issues, the video is made public. If manual monitoring is deemed necessary, the monitoring team is notified.
[1413] Step 13:
[1414] The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[1415] Step 14:
[1416] The server filters content that is inappropriate for the specified age and restricts playback on the device, preventing minors from accessing inappropriate content.
[1417] Example 2
[1418] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1419] In today's video streaming services, it is extremely difficult to efficiently process the large number of videos posted by users and provide relevant recommended videos to users. Furthermore, the task of appropriately filtering inappropriate content and ensuring user safety is a heavy burden. Furthermore, there is a demand for sophisticated recommendations and ad delivery based on user sentiment, and a system that can realize these elements in an integrated manner is needed.
[1420] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for uploading videos from a terminal, a means for checking the quality of the received videos and generating metadata, a means for analyzing the content of the videos and generating summaries, a means for categorizing the videos based on the summaries generated by the server and assigning tags and keywords, a means for adding the videos to a recommendation list based on user characteristics, a means for improving the video recommendation function using an emotion engine, a means for selecting and delivering optimal advertisements based on user characteristics, a means for determining whether automatic or manual monitoring is required based on the content of the videos, and a means for applying an age filter based on the estimated age of the user. This makes it possible to efficiently process videos posted by users, provide highly relevant recommendations, deliver optimized advertisements, and provide a safer usage environment.
[1421] "Terminal" is a general term for electronic devices that users use to upload video files and their meta information to a server, and includes, for example, smartphones and personal computers.
[1422] "Server" refers to a central computer system that receives videos uploaded by users and performs processes such as analysis, conversion, summary generation, recommendation and advertisement delivery.
[1423] "Uploading a video" refers to the process in which a user sends a video file from their own device to a server.
[1424] "Quality confirmation" is the process in which the server checks the quality of the video file received, including the format, resolution, and playback time, and converts it into a standard format.
[1425] "Metadata generation" refers to the process by which the server generates additional information, such as file size, playback time, and resolution, based on the received video file.
[1426] "Video content analysis" is the process in which the server uses AI technology to analyze video frames and audio to identify key scenes and important objects.
[1427] "Summary generation" refers to the process in which the server extracts key scenes and important information based on the analysis of the video content and generates a compact summary.
[1428] "Categorization" is the process of classifying videos into specific categories based on server-generated summaries and assigning tags and keywords.
[1429] A "recommendation list" is a list of videos that the server determines to be optimal based on the user's characteristics and provides them to the user.
[1430] An "emotion engine" refers to software or a system that has the function of analyzing emotions from a user's facial expressions, tone of voice, viewing history, etc.
[1431] "Advertisement selection and delivery" is the process in which the server selects the most appropriate advertisement based on user characteristics and sentiment analysis results, and delivers it to the user within the video at the appropriate time.
[1432] "Automatic monitoring" is a function in which the server analyzes the video content and automatically detects and evaluates inappropriate content based on predetermined criteria.
[1433] "Human monitoring" is a process in which a team of human monitors conducts additional verification on videos that are deemed high risk by automated monitoring.
[1434] "Age filter" refers to a feature that restricts content that is displayed or played based on the estimated age of the user.
[1435] This invention is a system that efficiently processes videos posted by users and delivers highly relevant recommendations and optimized advertisements. This system is mainly built based on communications between terminals, a server, and users, and specific embodiments are described below.
[1436] Video Ingestion and Analysis
[1437] 1. The device sends the video uploaded by the user to the server.
[1438] Users send video files and their meta information (title, description, etc.) to the server from their devices such as smartphones or PCs.
[1439] 2. The server receives the video, checks the quality, and generates metadata.
[1440] The server checks the format of the video file, converts it to a standard format if necessary, checks the video quality, and generates metadata (file size, duration, etc.).
[1441] 3. The server uses AI technology to analyze the video content and generate a summary.
[1442] The server analyzes the video frames to identify key scenes and important objects, and performs audio analysis to identify speakers and extract important phrases. Based on this information, a video summary is generated.
[1443] Categorizing the video
[1444] 1. The server categorizes the video based on the generated summary.
[1445] The server classifies videos into specific categories based on tags and keywords extracted from the summaries, and also assigns these tags and keywords to the video data.
[1446] 2. The server creates a recommendation list appropriate for each user.
[1447] The server selects the most suitable videos based on the user's viewing history and interests, and generates a recommendation list, which is then sent to the user's device.
[1448] Emotion analysis and recommendations using an emotion engine
[1449] 1. The server analyzes the user's emotions using an emotion engine.
[1450] The server analyzes the user's emotions based on facial expressions, tone of voice, viewing history, etc.
[1451] 2. The server reflects the results of the sentiment analysis in its video recommendations.
[1452] The server takes the emotional data into account and updates the recommendation list to provide more of the content the user enjoys.
[1453] Video Ad Optimization
[1454] 1. The server selects video ads based on user characteristics.
[1455] The server selects the most suitable advertisement based on the user's age, gender, interests, and emotional analysis results, and also sets the advertisement's meta information and playback timing.
[1456] 2. The server delivers ads within the video.
[1457] The server evaluates the relevance of the video content and inserts ads at the appropriate time, so that relevant ads are displayed when the user watches the video.
[1458] Determining what to monitor
[1459] 1. The server determines whether to monitor the video based on the summary.
[1460] The server detects whether the video summary contains inappropriate content (e.g., violence, adult content, etc.), performs a risk assessment, and determines whether manual supervision is necessary.
[1461] 2. The server will tag the data for manual monitoring as needed.
[1462] Videos that are deemed safe by automatic server monitoring will not be subject to manual monitoring. If manual monitoring is deemed necessary, the monitoring team will be notified.
[1463] Applying an age filter
[1464] 1. The server applies an age filter based on the user's estimated age.
[1465] The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[1466] 2. The server filters inappropriate content.
[1467] The server prevents content that is not appropriate for the specified age from being played on the user's device, thereby preventing minors from accessing inappropriate content.
[1468] Specific examples
[1469] Example 1: User-Submitted Pet Videos
[1470] 1. The user (device) uploads a video of their dog playing at home.
[1471] 2. The server receives the video, performs quality checks, and generates metadata.
[1472] 3. The server analyzes the video and generates keywords such as "pet," "dog," and "play."
[1473] 4. The server uses these keywords to recommend relevant pet videos to the user.
[1474] 5. The server performs sentiment analysis on the user, detects whether the user is happy, and recommends more pet videos based on that.
[1475] Example 2: Fashion-related video ads
[1476] 1. The server identifies that the user has fashion-related interests.
[1477] 2. The server uses an emotion engine to confirm that the user has positive emotions toward the fashion content.
[1478] 3. The server selects fashion-related video ads and inserts them into the video at the appropriate time.
[1479] 4. The device displays relevant ads while you watch fashion videos.
[1480] Example 3: Reducing monitoring costs
[1481] 1. The server verifies the summary of the posted video for inappropriate content.
[1482] 2. If the server determines that there is no problem, the video will be automatically made public and will no longer be subject to manual monitoring.
[1483] Prompt Sentence Examples
[1484] 1. Analyze the video content and extract key scenes and important phrases.
[1485] 2. Generate a recommendation list based on the user's viewing history.
[1486] 3. Use a sentiment engine to analyze user sentiment and provide appropriate content.
[1487] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1488] Step 1:
[1489] The device sends the video file and its meta information (title, description, etc.) to the server.
[1490] Input: A user uploads a video file and its meta information from their device.
[1491] Processing: The device sends the video file to the server.
[1492] Output: The video file and associated meta information are saved on the server.
[1493] Step 2:
[1494] The server checks the format of the video file received and converts it to a standard format if necessary.
[1495] Input: The submitted video file.
[1496] Processing: The server analyzes the format of the video file and converts it to a standard format (e.g., adjusting the resolution, converting the file format).
[1497] Output: Video files converted to standard format.
[1498] Step 3:
[1499] The server checks the quality of the received video and generates metadata.
[1500] Input: Video files converted to standard formats.
[1501] Processing: The server checks the quality of the video (e.g., resolution, duration, file size, etc.) and generates metadata (e.g., resolution, duration, file size, frame rate).
[1502] Output: Quality-checked video files and their metadata.
[1503] Step 4:
[1504] The server uses AI technology to analyze the video content and generate a summary.
[1505] Input: Quality-checked video files and metadata.
[1506] Processing: The server uses generative AI models to analyze video frames and audio, extract key scenes and phrases, identify speakers in the video, and generate a summary.
[1507] Output: Summary of video content and analysis results.
[1508] Step 5:
[1509] The server categorizes the video based on the generated summary and assigns tags and keywords.
[1510] Input: Summary of video content and analysis results.
[1511] Processing: Based on the summary, the server assigns tags and keywords (e.g., "pets," "dogs," "play," etc.) to the video and classifies it into a specific category.
[1512] Output: Video data with tags and keywords.
[1513] Step 6:
[1514] The server creates a recommendation list suitable for each user.
[1515] Input: Video data with tags and keywords, user viewing history.
[1516] Processing: The server selects relevant videos based on the user's viewing history and interests and generates a recommendation list.
[1517] Output: A list of recommendations suitable for the user.
[1518] Step 7:
[1519] The server analyzes the user's emotions using an emotion engine.
[1520] Input: User's viewing history, facial expressions, and tone of voice.
[1521] Processing: The server uses an emotion engine to analyze the user's facial expressions and tone of voice and extract emotional data.
[1522] Output: User sentiment analysis results.
[1523] Step 8:
[1524] The server reflects the results of the sentiment analysis in its video recommendations.
[1525] Input: User sentiment analysis results, recommendation list.
[1526] Processing: The server takes into account the emotion data and updates the recommendation list to provide more content that the user enjoys.
[1527] Output: An updated recommendation list reflecting the sentiment data.
[1528] Step 9:
[1529] The server selects video advertisements based on user characteristics.
[1530] Input: User's age, gender, interests, and sentiment analysis results.
[1531] Processing: The server selects the most suitable advertisement based on the user's characteristics and sets the advertisement's meta information and playback timing.
[1532] Output: Advertising data based on user characteristics.
[1533] Step 10:
[1534] The server delivers ads within the video.
[1535] Input: Advertising data based on user characteristics, video being watched.
[1536] Processing: The server evaluates the relevance of the video content and inserts advertisements at the appropriate time.
[1537] Output: Ads inserted during viewing.
[1538] Step 11:
[1539] The server determines whether the video should be monitored based on the summary.
[1540] Input: A summary of the video content.
[1541] Processing: The server detects inappropriate content (e.g., violence, adult content) from the video summary and performs a risk assessment.
[1542] Output: The result of automated or manual monitoring.
[1543] Step 12:
[1544] The server will tag the monitors manually as needed.
[1545] Input: Verdict of automated or manual monitoring.
[1546] Processing: The server automatically monitors and publishes videos that it determines to be problem-free, and notifies the monitoring team if it determines that there is a problem.
[1547] Output: published video or notification to surveillance team.
[1548] Step 13:
[1549] The server applies an age filter based on the user's estimated age.
[1550] Input: User registration information, activity history, video summary.
[1551] Processing: The server estimates the user's age and detects age-restricted keywords (e.g., "violence" and "adult") from the video summary.
[1552] Output: The filtered content.
[1553] Step 14:
[1554] The server filters inappropriate content.
[1555] Input: Content after age filter applied.
[1556] Processing: The server filters content that is inappropriate for the specified age group and prevents it from being played on the user's device.
[1557] Output: Content that is restricted from being accessed by minors.
[1558] (Application example 2)
[1559] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1560] In modern content distribution services, the large number of users posting videos necessitates efficient video summarization and the creation of recommendation lists tailored to each user. It is also important to optimize ad delivery, reduce monitoring costs, and provide content based on user characteristics. However, no system exists that can simultaneously address all of these challenges. In particular, video recommendations and ad delivery that take user emotions into account have yet to be realized, so an effective solution is needed.
[1561] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1562] In this invention, the server includes: a means for uploading videos from a terminal; a means for checking the quality of the received videos and generating metadata; a means for analyzing the video content and generating a summary; a means for categorizing the videos based on the generated summary and assigning tags and keywords; a means for adding the videos to a recommendation list based on user characteristics; a means for selecting and assigning optimal advertisements based on user characteristics; a means for determining whether automatic or manual monitoring is required based on the video content; a means for applying an age filter based on the estimated age of the user; a means for analyzing the user's emotions using emotion recognition technology and updating the video recommendation list based on the analysis; and a means for inserting video advertisements related to specific video categories into videos at appropriate times. This enables efficient creation of video summaries and recommendation lists, optimization of advertisement distribution, reduction of monitoring costs, and provision of advanced content that takes user emotions into consideration.
[1563] A "terminal" is an electronic device such as a smartphone or computer that a user uses to upload videos.
[1564] "Server" refers to a computer system on a network that receives videos, analyzes them, generates summaries, categorizes them, creates recommendation lists, selects advertisements, determines whether they are monitored automatically or manually, applies age filters, analyzes emotions, and delivers advertisements.
[1565] "Quality check" is a process in which the server checks the format, image quality, sound quality, etc. of the uploaded video and converts it to a standard format if necessary.
[1566] "Metadata generation" is the process by which the server generates the video's file size, playback time, and other related information.
[1567] "Summary generation" is a process in which the server analyzes the content of a video and creates a short summary of the video based on key scenes, important objects, and audio.
[1568] "Categorization" is a process in which the server classifies videos into specific categories based on the summary information generated and assigns related tags and keywords.
[1569] A "recommendation list" is a list in which the server selects specific videos based on the user's viewing history and interests and provides them to the user.
[1570] "Advertisement selection" is a process in which the server selects the most suitable advertisement based on the user's age, gender, interests, and emotional analysis results, and sets its meta information and playback timing.
[1571] "Automatic monitoring" is a process in which the server detects whether a video contains inappropriate content based on its summary information and determines whether manual monitoring is necessary.
[1572] "Age filter" is a process in which the server controls the viewing of inappropriate content based on the estimated age of the user.
[1573] "Emotion recognition technology" is a technology that allows the server to analyze the user's facial expressions, tone of voice, etc. to understand the user's emotional state.
[1574] A "video category" is a specific genre or theme that the server uses to categorize videos, such as pets, fashion, cooking, etc.
[1575] The system for implementing this invention is configured by linking user terminals, a server, and a network, and enables efficient video summarization, improved recommendation functionality, optimized ad distribution, reduced monitoring costs, and a safe usage environment.
[1576] First, a user uploads a video from a device such as a smartphone or PC. This device is used to send the video taken by the user to the server.
[1577] The server temporarily stores the received video and checks its quality. Specifically, it checks the video format and converts it to a standard format if necessary. It also checks the video's image quality and sound quality and generates metadata (such as file size and playback time). This process uses video editing libraries (e.g., OpenCV) and audio analysis libraries.
[1578] The server then uses AI technology to analyze the video content. This analysis involves identifying important scenes in each frame of the video and extracting key phrases from the audio data. The AI technology used here utilizes computer vision for video analysis and a speech recognition engine (e.g., Google Speech-to-Text API) for audio analysis. The generated summary is then compiled in a format that is easy for the user to view.
[1579] The server then categorizes the video based on the generated summary and assigns relevant tags and keywords to it, using natural language processing (NLP) techniques based on information extracted from the summary.
[1580] The server then generates a recommendation list based on the user's characteristics. The most suitable videos are selected based on the user's viewing history and interests, and this list is sent to the user's device. The recommendation engine uses machine learning models to analyze user behavior data.
[1581] Furthermore, the server uses emotion recognition technology to analyze the user's emotions. This analysis involves analyzing the user's facial expressions and tone of voice. The emotion engine uses, for example, the FER (Facial Emotion Recognition) library. Based on the analysis results, the video recommendation list according to the emotion is updated.
[1582] The server selects the most suitable advertisement based on the user's characteristics and inserts it into the video at the appropriate time. Ad selection takes into account age, gender, interests, and emotional analysis results. Ad delivery also uses an algorithm that evaluates the relevance of the advertisement to the video content.
[1583] Finally, the server has a means to determine whether to monitor automatically or manually based on the video content. The automatic monitoring system analyzes the summary of the uploaded video to detect inappropriate content, such as illegal content. If necessary, it notifies the manual monitoring team.
[1584] A concrete example is the automatic recommendation of pet videos. When a user uploads a video of their dog playing at home, the system receives and analyzes the video, generating tags such as "pet," "dog," and "play." Based on the user's past viewing history and sentiment analysis data, the system recommends new pet videos.
[1585] Example of an input prompt for a generative AI model:
[1586] "Enter the video title, description, and uploaded file path. Example: Title: 'Dogs playing at home', Description: 'Dogs playing filmed at home', File path: 'path / to / video.mp4'"
[1587] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1588] Step 1:
[1589] Videos are uploaded from the device
[1590] A user shoots a video using a smartphone or computer and accesses the upload page. The prompt message is "Please enter the video title, description, and uploaded file path. Example: Title: 'Dog playing at home', Description: 'Dog playing filmed at home', File path: 'path / to / video.mp4'." When the user enters the title, description, and video file and presses the upload button, the video file and metadata are sent to the server. The input is the video file and its metadata specified by the user, and the output is the video file and its metadata stored on the server.
[1591] Step 2:
[1592] The server checks the quality of the received video and generates metadata.
[1593] The server temporarily stores the received video file and checks its format, image quality, and sound quality. For example, it uses OpenCV to check the video resolution and pixel quality and converts it to a standard format if necessary. It also uses an audio analysis library to check the sound quality. Metadata such as file size, playback time, and resolution are generated. The input is the received video file, and the output is the generated metadata.
[1594] Step 3:
[1595] The server analyzes the video content and generates a summary
[1596] The server uses AI technology to analyze the video content. Each frame of the video is analyzed using computer vision technology to identify key scenes. For example, OpenCV and deep learning models are used to detect key objects and movements. The audio is converted to text using the Google Speech-to-Text API, and key phrases are extracted. A summary is generated based on this information. The input is each frame of the video and audio data, and the output is a video summary.
[1597] Step 4:
[1598] The server categorizes the video based on the generated summary and assigns tags and keywords.
[1599] The server classifies videos into specific categories based on information obtained from the video summaries and assigns associated tags and keywords. For example, natural language processing techniques are used to extract keywords obtained from the summaries and map them to predefined categories. The input is the generated video summaries, and the output is the classified video categories and assigned tags and keywords.
[1600] Step 5:
[1601] The server adds videos to the recommendation list based on the user's characteristics.
[1602] The server analyzes the user's viewing history, interests, and past viewing data to select the most suitable videos and generate a recommendation list. Here, a machine learning model analyzes the user's behavioral data and lists highly relevant videos. For example, a collaborative filtering algorithm is used. The input is the user's viewing history and interest data, and the output is a recommendation list.
[1603] Step 6:
[1604] The server uses emotion recognition technology to analyze the user's emotions and updates the video recommendation list based on that.
[1605] The server performs facial expression analysis and voice tone analysis to recognize the user's emotions. For example, it uses the FER (Facial Emotion Recognition) library to analyze facial expressions in frames and identify the user's emotions. Based on the results, it adds more types of videos that the user enjoys to the recommendation list. The input is each frame of the video and audio data, and the output is an updated recommendation list.
[1606] Step 7:
[1607] The server selects the most suitable advertisement based on the user's characteristics and inserts it into the video at the appropriate time.
[1608] The server selects the most suitable advertisement based on the user's age, gender, interests, and emotional analysis results, and determines the timing of its insertion into the video. For example, collaborative filtering and emotional analysis data can be combined to identify the user's areas of interest, select highly relevant advertisements, and incorporate them into the video frame. The input is the user's characteristics and the results of the emotional analysis, and the output is a video containing the inserted advertisement.
[1609] Step 8:
[1610] The server determines whether to monitor automatically or manually based on the video content.
[1611] The server uses the generated summary to check for inappropriate content in the video using an automated monitoring system. For example, it uses content analysis algorithms to detect problematic scenes or audio, performs a risk assessment, and notifies a human monitoring team if necessary. The input is the generated video summary, and the output is the monitoring decision.
[1612] Step 9:
[1613] The server applies an age filter based on the user's estimated age.
[1614] The server estimates the user's age based on their registration information and activity history, and restricts access to inappropriate content. For example, it extracts age-related keywords from the summary information and filters videos that do not match the user's age. The input is the user's estimated age and the video summary, and the output is a filtered list of videos.
[1615] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1616] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1617] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1618] [Fourth embodiment]
[1619] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1620] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1621] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1622] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1623] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1624] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1625] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1626] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1627] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1628] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1629] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1630] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1631] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1632] This invention is a system that efficiently summarizes videos posted by users and, based on the summaries, enables recommendation functions, advertisement distribution, reduced monitoring costs, and a safe usage environment. This system is built on communications between terminals, servers, and users. Below, the processing of the system's program is explained in natural language, with specific examples.
[1633] Video Ingestion and Analysis
[1634] 1. The device sends the video uploaded by the user to the server.
[1635] Device: The user uploads the video file and its meta information (title, description, etc.) to the server from a device such as a smartphone or PC.
[1636] 2. The server receives the video and performs the ingest process.
[1637] Server: Checks the format of the video file and converts it to a standard format if necessary. It also checks the video quality and generates metadata (file size, duration, etc.).
[1638] 3. The server uses AI technology to analyze the video content and generate a summary.
[1639] Server: Analyzes video frames to identify key scenes and important objects. Also, analyzes audio to identify speakers and extract important phrases. Then, it generates a video summary based on this information.
[1640] Categorizing the video
[1641] 1. The server categorizes the video based on the generated summary.
[1642] Server: Classifies videos into specific categories based on tags and keywords extracted from the summaries, and assigns these tags and keywords to the video data.
[1643] 2. The server creates a recommendation list suitable for each user.
[1644] Server: Based on the user's viewing history and interests, the server selects the most suitable videos and generates a recommendation list. This list is sent to the user's device.
[1645] Video Ad Optimization
[1646] 1. The server selects video ads based on user characteristics.
[1647] Server: Selects the most suitable advertisement based on the user's characteristics such as age, gender, and interests. Also sets the advertisement meta information and playback timing.
[1648] 2. The server delivers ads within the video.
[1649] Server: Evaluates the relevance of the video content and inserts ads at the appropriate time. When a user watches a video, relevant ads are displayed.
[1650] Determining what to monitor
[1651] 1. The server determines whether to monitor the video based on the summary.
[1652] Server: Detects whether the video summary contains inappropriate content (e.g. violence, adult content, etc.), performs a risk assessment, and determines whether manual supervision is required.
[1653] 2. The server will tag the data for manual monitoring as needed.
[1654] Server: Videos that are deemed safe by automatic monitoring are exempt from manual monitoring. If manual monitoring is deemed necessary, the monitoring team is notified.
[1655] Applying an age filter
[1656] 1. The server applies an age filter based on the user's estimated age.
[1657] Server: Estimates the user's age from their registration information and activity history. Detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[1658] 2. The server filters inappropriate content.
[1659] Server: Content that is not appropriate for the specified age group will not be played on the user's device, thereby preventing minors from accessing inappropriate content.
[1660] Specific examples
[1661] Example 1: User-Submitted Pet Videos
[1662] 1. The user (device) uploads a video of their dog playing at home.
[1663] 2. The server receives the video, performs quality checks, and generates metadata.
[1664] 3. The server analyzes the video and generates keywords such as "pet," "dog," and "play."
[1665] 4. The server uses these keywords to recommend pet videos relevant to the user.
[1666] Example 2: Fashion-related video ads
[1667] 1. The server identifies that the user has fashion-related interests.
[1668] 2. The server selects fashion-related video ads and inserts them into the video at the appropriate time.
[1669] 3. The device displays relevant ads while you watch fashion videos.
[1670] Example 3: Reducing monitoring costs
[1671] 1. The server verifies the summary of the posted video for inappropriate content.
[1672] 2. If the server determines that there is no problem, the video will be automatically made public and will no longer be subject to manual monitoring.
[1673] This system enables the provision of content tailored to user characteristics, optimization of advertising revenue, and reduction of human monitoring costs.
[1674] The processing flow will be explained below.
[1675] Step 1:
[1676] A user uses a terminal to upload a video file and its meta information (title, description, etc.) to a server.
[1677] Step 2:
[1678] The server checks the format of the video file received and converts it to a standard format if necessary, checks the video quality, and generates metadata (file size, playback time, etc.).
[1679] Step 3:
[1680] The server begins analyzing the video content using AI technology, specifically by analyzing the frames of the video and identifying key scenes and important objects.
[1681] Step 4:
[1682] The server performs audio analysis to identify speakers and extract key phrases, and generates a video summary based on the results of this analysis.
[1683] Step 5:
[1684] Based on the generated summary, the server assigns tags and keywords to the video and classifies it into a specific category, thereby categorizing the video.
[1685] Step 6:
[1686] The server selects appropriate videos based on the user's viewing history and interests, and generates a recommendation list for the user. This list is sent to the device.
[1687] Step 7:
[1688] The server selects the most suitable video ad based on user characteristics (age, gender, interests, etc.) and also sets the ad's meta information and playback timing.
[1689] Step 8:
[1690] The server evaluates the relevance of the video content and inserts ads at the appropriate time, and displays relevant ads when the device is watching the video.
[1691] Step 9:
[1692] Based on the summary, the server determines whether the video is suitable for automatic monitoring or requires manual monitoring. Inappropriate content (e.g., violence, adult content, etc.) is detected.
[1693] Step 10:
[1694] The server performs a risk assessment and, if it determines that there are no problems with automated monitoring, it makes the video public. If it determines that manual monitoring is necessary, it notifies the monitoring team.
[1695] Step 11:
[1696] The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the summary.
[1697] Step 12:
[1698] The server filters content that is inappropriate for the specified age and prevents the video from being played on the device, thereby preventing minors from accessing inappropriate content.
[1699] Example 1
[1700] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1701] Modern video streaming services require efficient analysis and classification of the vast amount of videos uploaded by users, and the provision of appropriate recommendations and advertisements. However, manual video analysis and recommendation generation requires a great deal of effort and time, resulting in high operational costs, making efficient and automated processing methods necessary. Automation is also required for providing personalized advertisements based on user characteristics, filtering content suitable for underage users, and monitoring video content. An efficient and reliable system is needed to solve these challenges.
[1702] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1703] In this invention, the server includes a means for uploading videos from a terminal, a means for converting the format of the received video, checking the quality, and generating metadata, a means for analyzing the content of the video using a generation AI model and generating a summary, a means for categorizing the video based on the generated summary and assigning tags and keywords, a means for adding the video to a recommendation list based on user characteristics, a means for selecting and delivering optimal advertisements based on user characteristics, a means for determining whether automatic or manual monitoring is required based on the video content, and a means for applying an age filter based on the estimated age of the user. This enables efficient analysis of videos uploaded by users, provision of appropriate recommendations and advertisements, automated content monitoring, and age-appropriate filtering.
[1704] A "terminal" is a device used by a user, such as a smartphone or computer, used to upload videos.
[1705] The "server" is a central device that receives, analyzes, and processes video data over the network, and plays a central role in this system.
[1706] A "video" is a file containing a series of video and audio files uploaded by a user.
[1707] "Format conversion" is the process of converting received video files into a standard format.
[1708] "Quality check" is the process of checking the quality of the received video, such as its resolution and playback time.
[1709] "Metadata" refers to various information related to a video file (e.g., file size, playback time, resolution, etc.).
[1710] A "generative AI model" is an artificial intelligence technology used to analyze video content and generate summaries.
[1711] "Video analysis" is the process of analyzing video frames and audio to identify important scenes and objects.
[1712] "Summary" refers to a short summary of the main content extracted from the results of video analysis.
[1713] "Tags" refer to keywords or labels that are added to indicate the content of a video.
[1714] "Keywords" refer to words or short phrases related to the video content.
[1715] A "recommendation list" is a list of videos selected based on the user's characteristics and interests.
[1716] "Advertisement" refers to information including product promotions and service information, which is inserted into videos so that users can view them.
[1717] "Automatic monitoring" is the process of automatically checking video content to detect inappropriate content.
[1718] "Manual monitoring" is a manual monitoring process carried out by specialized staff when automatic monitoring is difficult to judge.
[1719] "Estimated age" is the age calculated based on the user's registration information and activity history.
[1720] An "age filter" is a mechanism that provides only appropriate content based on the user's estimated age.
[1721] This invention is a system that efficiently summarizes videos posted by users and provides recommendation functions, advertisement distribution, reduced monitoring costs, and a safe usage environment based on the summaries. This system is built on communications between terminals, a server, and users.
[1722] Video Ingestion and Analysis
[1723] When a user uploads a video to a server from a device, the user uses a device such as a smartphone or PC. Specifically, the user sends the video file and its meta information (title, description, etc.) to the server through an interface provided by the device.
[1724] The server receives video files uploaded by users. The received video files are first converted into a different format. Software such as ffmpeg is used for this process. Next, the quality of the video is checked and metadata (file size, resolution, playback time, etc.) is generated based on that quality.
[1725] The server uses a generative AI model to analyze the video content and generate a summary. Specifically, it uses an object detection model such as YOLO (You Only Look Once) to identify key scenes and important objects, and then performs audio analysis using the Google Cloud Speech-to-Text API. Then, based on this information, it uses a natural language generation model (e.g., GPT-3) to generate a video summary.
[1726] Video categorization and recommendations
[1727] The server categorizes the videos based on the generated summaries. The videos are classified into specific categories based on tags and keywords extracted from the summaries. The classified videos are then assigned tags and keywords to improve searchability and filtering capabilities.
[1728] The server then generates a recommendation list based on the user's characteristics. Taking into account the user's viewing history and interests, the server selects the most suitable videos to create the recommendation list. This recommendation list is then sent to the user's device, where the user can watch the recommended videos through an interface.
[1729] Ad optimization and delivery
[1730] The server selects the most suitable advertisement based on the video summary and user characteristics. Specifically, it selects advertisements based on the user's characteristics such as age, gender, and interests, and sets the meta information and playback timing. The server references a database provided by the advertiser to select advertisements.
[1731] The server delivers advertisements at appropriate times within the video. For example, while watching a video about pets, it displays an advertisement for related pet food. This can be expected to have a high advertising effect.
[1732] Determining who should be monitored and applying age filters
[1733] The server determines whether a video should be subject to automated or manual monitoring based on the summary. The video summary is used to detect whether it contains inappropriate content and perform a risk assessment. Based on this, the process of determining whether to subject the video to automated or manual monitoring uses natural language processing technology.
[1734] Furthermore, the server applies an age filter based on the user's estimated age. The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the video summary. Inappropriate content is filtered so that it cannot be played on devices belonging to underage users. Based on these filtering rules, only appropriate videos are provided.
[1735] Specific examples
[1736] Example 1: User-Submitted Pet Videos
[1737] A user uploads a video of their dog playing at home. The server receives the video, checks the quality, and generates metadata. The server then analyzes the video and generates keywords such as "pet," "dog," and "play." The server then recommends relevant pet videos to the user based on these keywords.
[1738] Example 2: Fashion-related video ads
[1739] The server identifies the user's interest in fashion. The server then selects fashion-related video ads and inserts them into the video at the appropriate time. The relevant ads are then displayed on the user's device while they are watching the fashion video.
[1740] Example 3: Reducing monitoring costs
[1741] The server verifies the summary of the posted video to see if it contains inappropriate content. If it determines there are no problems, it automatically publishes the video and removes it from the scope of human monitoring.
[1742] Prompt Sentence Examples
[1743] "Recommend other videos related to this pet video."
[1744] "Insert fashion-related ads at the best possible time."
[1745] Please determine if this video is inappropriate.
[1746] This system makes it possible to provide content tailored to user characteristics, optimize advertising revenue, and reduce manual monitoring costs.
[1747] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1748] Step 1:
[1749] The device sends the video uploaded by the user to the server. The user clicks the upload button through the device interface, selects the video file (e.g., mp4, avi) and its meta information (title, description), and sends it to the server.
[1750] Input: Video file and meta information
[1751] Output: Video file and meta information sent to the server
[1752] Step 2:
[1753] The server receives the video, converts the format, checks the quality, and generates metadata. The server checks the format of the received video file and converts it to a standard format (e.g., mp4). This process uses tools such as ffmpeg. It also checks the quality of the video, such as resolution, playback time, and file size, and generates this information as metadata.
[1754] Input: Video file and meta information sent to the server
[1755] Output: Format converted video file and generated metadata
[1756] Step 3:
[1757] The server uses a generative AI model to analyze the video content and generate a summary. Specifically, it uses an object detection model such as YOLO to analyze video frames and identify key scenes and important objects. It also performs audio analysis using the Google Cloud Speech-to-Text API to identify speakers and extract key phrases. It then uses this data to generate a video summary using a natural language generation model (e.g., GPT-3).
[1758] Input: Format converted video file and metadata
[1759] Output: Video summary
[1760] Step 4:
[1761] The server categorizes the video based on the generated summary and assigns tags and keywords. Based on the information obtained from the summary, the server classifies the video into an appropriate category (e.g., pets, sports, entertainment), and assigns tags and keywords related to that category to the video data.
[1762] Input: Video summary
[1763] Output: Categorized videos with tags and keywords
[1764] Step 5:
[1765] The server adds videos to a recommendation list based on user characteristics. The server analyzes the user's viewing history and interests, selects the most suitable videos, and generates a recommendation list. This list is sent to the user's device.
[1766] Input: User's viewing history and interests, as well as categorized videos, tags, and keywords
[1767] Output: Recommendation list based on user characteristics
[1768] Step 6:
[1769] The server selects and delivers the most suitable advertisement based on the user's characteristics. The server selects the appropriate advertisement based on the user's characteristics information such as age, gender, and interests. Specifically, it retrieves the advertisement that best suits the user's characteristics from the advertisement database and sets its meta information and playback timing.
[1770] Input: User characteristics (age, gender, interests, etc.) and advertising database
[1771] Output: Optimal advertisement based on user characteristics
[1772] Step 7:
[1773] The server determines whether the video should be monitored automatically or manually based on its content. The server then checks the video summary for inappropriate content and assesses the content risk. This process uses natural language processing technology.
[1774] Input: Video summary
[1775] Output: Judgment results of automatic or manual monitoring
[1776] Step 8:
[1777] The server applies an age filter based on the user's estimated age. The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords from the video summary. Inappropriate content is filtered so that it cannot be played on devices belonging to minors.
[1778] Input: User registration information, activity history, video summary
[1779] Output: Age-filtered video list
[1780] (Application example 1)
[1781] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1782] In recent years, with the proliferation of video content, a huge amount of videos has been uploaded to online platforms. However, there is a lack of systems that can efficiently summarize videos and recommend appropriate content to users. Furthermore, manually monitoring video content and filtering inappropriate content is extremely time-consuming and costly. Furthermore, optimization of advertisements based on user characteristics is insufficient, and efficient advertisement delivery is required. To address these issues, a system is needed that automates video analysis and recommendation, advertisement optimization, and content monitoring, while providing a safe viewing environment.
[1783] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1784] In this invention, the server includes a means for uploading videos from a terminal, a means for checking the quality of the received videos and generating metadata, a means for analyzing the video content and generating a summary, a means for categorizing the videos based on the generated summary and assigning tags and keywords, a means for adding the videos to a recommendation list based on user characteristics, a means for selecting and assigning optimal advertisements based on user characteristics, a means for determining whether automatic or manual monitoring is required based on the video content, a means for applying an age filter based on the estimated age of the user, a means for saving extracted frame images and transcribing and analyzing audio, a means for providing a function for recommending similar content based on responses from the server, and a means for inserting advertisements and optimizing based on user characteristics as part of a content distribution service. This enables efficient video summarization and recommendation functions, optimal advertisement distribution, reduced monitoring costs, and a safe usage environment.
[1785] - "Device" refers to the device to which the video is uploaded, such as a smartphone or computer.
[1786] "Server" refers to the computing system that processes the received video, checks the quality, generates metadata, and analyzes the video content.
[1787] "Video content analysis" refers to the process of analyzing video frames and audio data to extract summaries, important scenes, and objects.
[1788] "Generating a summary" refers to extracting the main points from the analyzed video data and summarizing them concisely.
[1789] "Categorizing videos" refers to classifying videos into specific categories based on the generated summary and assigning relevant tags and keywords.
[1790] "Add to recommendation list" refers to selecting suitable videos based on the user's viewing history and interests and adding them to the display list.
[1791] "Selecting and delivering ads" refers to selecting the most suitable ads based on user characteristics and inserting them while the video is playing.
[1792] "Determine whether to use automated or manual monitoring" refers to determining whether a video contains inappropriate content based on a summary of the video and deciding whether to use automated or, if necessary, manual monitoring.
[1793] "Applying age filters" refers to restricting the content available to a user based on their estimated age.
[1794] "Save frame images" means saving still images extracted from a video, usually at 1-second intervals.
[1795] "Audio transcription and analysis" refers to the process of converting video audio into text data and extracting key phrases and keywords.
[1796] Providing a "similar content recommendation function" refers to a function that automatically suggests relevant videos based on a user's viewing history and characteristics.
[1797] "Ad insertion and optimization based on user characteristics" refers to inserting ads into videos at the appropriate time and providing the optimal advertising plan based on the user's characteristics.
[1798] This invention provides a system that efficiently summarizes videos posted by users and, based on the summaries, enables recommendation functions, advertisement distribution, reduced monitoring costs, and a safe usage environment. Detailed modes for carrying out the invention are described below.
[1799] Uploading videos from your device
[1800] Users upload video files to the server using devices such as smartphones or PCs. At this time, the video file and its meta information (title, description, etc.) are also sent.
[1801] Server receives and processes video
[1802] The server checks the quality of the received video files, converts them to a standard format if necessary, and generates metadata (file size, playback time, etc.) for the video.
[1803] Video content analysis and summary generation
[1804] The server analyzes the video frames to identify key scenes and important objects, and performs audio analysis to identify speakers and extract key phrases. Based on this information, a video summary is generated. This process includes saving the captured frame images and transcribing and analyzing the audio extracted from the video.
[1805] Video categorization and recommendations
[1806] The server categorizes the videos based on the generated summaries and assigns tags and keywords. Based on the user's characteristics, it selects the most suitable videos and generates a recommendation list. Based on the response from the server, a recommendation function for similar content is provided.
[1807] Ad optimization and insertion
[1808] The server selects the most suitable advertisement based on the user's characteristics and inserts it at a timing relevant to the video content. As part of the content distribution service, it performs advertisement insertion and optimization based on user characteristics.
[1809] Video monitoring and age filtering
[1810] The server determines whether to monitor the content based on the video summary, automatically or manually, and applies an age filter based on the user's estimated age to filter out inappropriate content.
[1811] Hardware and software used
[1812] The following hardware and software are used to implement this system.
[1813] Hardware: Smartphones, PCs
[1814] Software: MoviePy (video editing library), Vosk (voice recognition library), OpenCV (image processing library)
[1815] Specific examples
[1816] Example 1: User-Submitted Pet Videos
[1817] When a user uploads a video of their dog playing at home, the server receives the video, checks the quality, and generates metadata. It then analyzes the video to generate keywords such as "pet," "dog," and "play," and recommends related pet videos.
[1818] Example 2: Fashion-related video ads
[1819] The server identifies that the user is interested in fashion, selects fashion-related video advertisements, and inserts them into the videos at the appropriate time. The relevant advertisements are displayed while the user is watching the fashion video.
[1820] Prompt Sentence Examples
[1821] text
[1822] Analyzes the user-uploaded video "example_video.mp4" and generates a summary. Then, recommends similar videos and inserts relevant ads. Finally, applies supervision and age filters.
[1823] In this way, the present invention makes it possible to provide efficient video summarization and recommendation functions, optimal advertisement distribution, reduced monitoring costs, and a safe usage environment.
[1824] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1825] Step 1:
[1826] A user uses a terminal to upload a video file to a server.
[1827] Input: Video file, meta information (title, description, etc.)
[1828] Output: Video files and meta information are saved on the server
[1829] Specific operation: A user selects a video file and presses the upload button on their smartphone or PC. The device then sends the selected video file and its meta information to the server as an HTTP request.
[1830] Step 2:
[1831] The server checks the quality of the received video file and generates metadata.
[1832] Input: Uploaded video file, meta information
[1833] Output: Video quality evaluation results, metadata (file size, playing time, etc.)
[1834] What it does: The server checks the format of the video file and converts it to a standard format if necessary. It also evaluates the video quality and generates metadata such as file size, duration, and resolution.
[1835] Step 3:
[1836] The server performs frame analysis and audio analysis of the video content and generates a summary.
[1837] Input: Video file
[1838] Output: Video summary, list of important scenes and objects, audio transcription
[1839] Specific operations: Extract key scenes from each frame of a video using MoviePy and OpenCV. Transcribe audio data using Vosk and extract key phrases and keywords. Combine these to generate a video summary.
[1840] Step 4:
[1841] The server categorizes the video based on the generated summary and assigns tags and keywords.
[1842] Input: Video summary
[1843] Output: Video data with tags and keywords
[1844] How it works: The server classifies videos into specific categories based on phrases and keywords extracted from the summaries, automatically assigning tags such as "pets," "dogs," and "play."
[1845] Step 5:
[1846] The server generates a recommendation list based on user characteristics.
[1847] Input: User viewing history, interests, summary data
[1848] Output: Recommendation list for each user
[1849] How it works: The server analyzes the user's past viewing history and interests (e.g., pet videos) and selects relevant videos from the summary data, thereby generating a customized recommendation list for each user.
[1850] Step 6:
[1851] The server selects the most suitable advertisement based on user characteristics and inserts it at the appropriate time within the video.
[1852] Input: User characteristics, video summary, relevant advertising data
[1853] Output: Video with ads
[1854] Specific operation: The server selects the most suitable advertisement based on the user's characteristics such as age, gender, and interests, and inserts the selected advertisement at the relevant timing in the video (e.g., scene change).
[1855] Step 7:
[1856] The server determines which videos to monitor based on the video summary.
[1857] Input: Video summary
[1858] Output: Judgment result of whether the target is automatic or manual monitoring
[1859] How it works: The server evaluates the summary and determines whether it contains inappropriate content. If the automated monitoring finds no problems, the video is left public. If manual monitoring is deemed necessary, the monitoring team is notified.
[1860] Step 8:
[1861] The server applies an age filter based on the user's estimated age.
[1862] Input: User registration information, activity history, video summary
[1863] Output: Filtered video content
[1864] Specific operation: The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords from the video summary. It then filters the video to prevent underage users from accessing inappropriate content.
[1865] Examples of prompt statements
[1866] Analyzes the user-uploaded video "example_video.mp4" and generates a summary. Then, recommends similar videos and inserts relevant ads. Finally, applies supervision and age filters.
[1867] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1868] This invention is a system that efficiently summarizes videos posted by users, improving recommendation functionality, optimizing ad delivery, reducing monitoring costs, and providing a safe usage environment. Furthermore, by combining it with an emotion engine that analyzes user emotions, more advanced recommendations and ad delivery become possible. This system is built on communications between terminals, servers, and users. Below, we explain the system's program processing in natural language, including specific examples.
[1869] Video Ingestion and Analysis
[1870] 1. The device sends the video uploaded by the user to the server.
[1871] Device: The user uploads the video file and its meta information (title, description, etc.) to the server from a device such as a smartphone or PC.
[1872] 2. The server receives the video and performs the ingest process.
[1873] Server: Checks the format of the video file and converts it to a standard format if necessary. It also checks the video quality and generates metadata (file size, duration, etc.).
[1874] 3. The server uses AI technology to analyze the video content and generate a summary.
[1875] Server: Analyzes video frames to identify key scenes and important objects. Also, analyzes audio to identify speakers and extract important phrases. Then, it generates a video summary based on this information.
[1876] Categorizing the video
[1877] 1. The server categorizes the video based on the generated summary.
[1878] Server: Classifies videos into specific categories based on tags and keywords extracted from the summaries, and assigns these tags and keywords to the video data.
[1879] 2. The server creates a recommendation list suitable for each user.
[1880] Server: Based on the user's viewing history and interests, the server selects the most suitable videos and generates a recommendation list. This list is sent to the user's device.
[1881] Emotion analysis and recommendations using an emotion engine
[1882] 1. The server uses an emotion engine to analyze the user's emotions.
[1883] Server: Analyzes emotions based on the user's facial expressions, tone of voice, viewing history, etc.
[1884] 2. The server reflects the results of the sentiment analysis in its video recommendations.
[1885] Server: Updates the recommendation list taking into account emotional data to provide more of the content the user enjoys.
[1886] Video Ad Optimization
[1887] 1. The server selects video ads based on user characteristics.
[1888] Server: Selects the most suitable advertisement based on the user's age, gender, interests, and emotional analysis results. Also sets the advertisement meta information and playback timing.
[1889] 2. The server delivers ads within the video.
[1890] Server: Evaluates the relevance of the video content and inserts ads at the appropriate time. When a user watches a video, relevant ads are displayed.
[1891] Determining what to monitor
[1892] 1. The server determines whether to monitor the video based on the summary.
[1893] Server: Detects whether the video summary contains inappropriate content (e.g. violence, adult content, etc.), performs a risk assessment, and determines whether manual supervision is required.
[1894] 2. The server will tag the data for manual monitoring as needed.
[1895] Server: Videos that are deemed safe by automatic monitoring are exempt from manual monitoring. If manual monitoring is deemed necessary, the monitoring team is notified.
[1896] Applying an age filter
[1897] 1. The server applies an age filter based on the user's estimated age.
[1898] Server: Estimates the user's age from their registration information and activity history. Detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[1899] 2. The server filters inappropriate content.
[1900] Server: Content that is not appropriate for the specified age group is not played on the user's device, thereby preventing minors from accessing inappropriate content.
[1901] Specific examples
[1902] Example 1: User-Submitted Pet Videos
[1903] 1. The user (device) uploads a video of their dog playing at home.
[1904] 2. The server receives the video, performs quality checks, and generates metadata.
[1905] 3. The server analyzes the video and generates keywords such as "pet," "dog," and "play."
[1906] 4. The server uses these keywords to recommend pet videos relevant to the user.
[1907] 5. The server analyzes the user's emotions, detects whether the user is happy, and recommends more pet videos based on that.
[1908] Example 2: Fashion-related video ads
[1909] 1. The server identifies that the user has fashion-related interests.
[1910] 2. The server uses the emotion engine to confirm that the user has positive emotions toward the fashion content.
[1911] 3. The server selects fashion-related video ads and inserts them into the video at the appropriate time.
[1912] 4. The device displays relevant ads while you watch fashion videos.
[1913] Example 3: Reducing monitoring costs
[1914] 1. The server verifies the summary of the posted video for inappropriate content.
[1915] 2. If the server determines that there is no problem, the video will be automatically made public and will no longer be subject to manual monitoring.
[1916] Note
[1917] This system utilizes an emotion engine to provide content tailored to user characteristics, optimize advertising revenue, and reduce manual monitoring costs. In addition, by taking into account the user's emotional state, an improved user experience is expected.
[1918] The processing flow will be explained below.
[1919] Step 1:
[1920] A user uses a terminal to upload a video file and its meta information (title, description, etc.) to a server.
[1921] Step 2:
[1922] The server checks the format of the received video file and converts it to a standard format if necessary, as well as checking the video quality and generating metadata (file size, duration, etc.).
[1923] Step 3:
[1924] The server begins analyzing the video content using AI technology, specifically by analyzing the frames of the video and identifying key scenes and important objects.
[1925] Step 4:
[1926] The server performs audio analysis to identify speakers and extract key phrases, and generates a summary of the video based on the results of this analysis.
[1927] Step 5:
[1928] Based on the generated summary, the server assigns tags and keywords to the video and classifies it into a specific category, thereby categorizing the video.
[1929] Step 6:
[1930] The server selects appropriate videos based on the user's viewing history and interests, and generates a recommendation list for the user. This list is sent to the device.
[1931] Step 7:
[1932] The server uses an emotion engine to analyze the user's emotions. Specifically, it collects data such as the user's facial expressions, voice, and viewing history to evaluate their emotional state.
[1933] Step 8:
[1934] The server updates the recommendation list based on the results of the sentiment analysis. Specifically, it prioritizes recommendations of content that users enjoy watching.
[1935] Step 9:
[1936] The server selects video ads based on user characteristics and sentiment analysis results, selects the most suitable ad, and also sets its meta information and playback timing.
[1937] Step 10:
[1938] The server evaluates the relevance of the video content and inserts advertisements at the appropriate time, and the device displays the advertisements.
[1939] Step 11:
[1940] The server automatically determines what to monitor based on the video summary, and detects inappropriate content (e.g., violence, adult content, etc.) from the summary.
[1941] Step 12:
[1942] The server performs a risk assessment and, if automated monitoring determines there are no issues, the video is made public. If manual monitoring is deemed necessary, the monitoring team is notified.
[1943] Step 13:
[1944] The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[1945] Step 14:
[1946] The server filters content that is inappropriate for the specified age and restricts playback on the device, preventing minors from accessing inappropriate content.
[1947] Example 2
[1948] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1949] In today's video streaming services, it is extremely difficult to efficiently process the large number of videos posted by users and provide relevant recommended videos to users. Furthermore, the task of appropriately filtering inappropriate content and ensuring user safety is a heavy burden. Furthermore, there is a demand for sophisticated recommendations and ad delivery based on user sentiment, and a system that can realize these elements in an integrated manner is needed.
[1950] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for uploading videos from a terminal, a means for checking the quality of the received videos and generating metadata, a means for analyzing the content of the videos and generating summaries, a means for categorizing the videos based on the summaries generated by the server and assigning tags and keywords, a means for adding the videos to a recommendation list based on user characteristics, a means for improving the video recommendation function using an emotion engine, a means for selecting and delivering optimal advertisements based on user characteristics, a means for determining whether automatic or manual monitoring is required based on the content of the videos, and a means for applying an age filter based on the estimated age of the user. This makes it possible to efficiently process videos posted by users, provide highly relevant recommendations, deliver optimized advertisements, and provide a safer usage environment.
[1951] "Terminal" is a general term for electronic devices that users use to upload video files and their meta information to a server, and includes, for example, smartphones and personal computers.
[1952] "Server" refers to a central computer system that receives videos uploaded by users and performs processes such as analysis, conversion, summary generation, recommendation and advertisement delivery.
[1953] "Uploading a video" refers to the process in which a user sends a video file from their own device to a server.
[1954] "Quality confirmation" is the process in which the server checks the quality of the video file received, including the format, resolution, and playback time, and converts it into a standard format.
[1955] "Metadata generation" refers to the process by which the server generates additional information, such as file size, playback time, and resolution, based on the received video file.
[1956] "Video content analysis" is the process in which the server uses AI technology to analyze video frames and audio to identify key scenes and important objects.
[1957] "Summary generation" refers to the process in which the server extracts key scenes and important information based on the analysis of the video content and generates a compact summary.
[1958] "Categorization" is the process of classifying videos into specific categories based on server-generated summaries and assigning tags and keywords.
[1959] A "recommendation list" is a list of videos that the server determines to be optimal based on the user's characteristics and provides them to the user.
[1960] An "emotion engine" refers to software or a system that has the function of analyzing emotions from a user's facial expressions, tone of voice, viewing history, etc.
[1961] "Advertisement selection and delivery" is the process in which the server selects the most appropriate advertisement based on user characteristics and sentiment analysis results, and delivers it to the user within the video at the appropriate time.
[1962] "Automatic monitoring" is a function in which the server analyzes the video content and automatically detects and evaluates inappropriate content based on predetermined criteria.
[1963] "Human monitoring" is a process in which a team of human monitors conducts additional verification on videos that are deemed high risk by automated monitoring.
[1964] "Age filter" refers to a feature that restricts content that is displayed or played based on the estimated age of the user.
[1965] This invention is a system that efficiently processes videos posted by users and delivers highly relevant recommendations and optimized advertisements. This system is mainly built based on communications between terminals, a server, and users, and specific embodiments are described below.
[1966] Video Ingestion and Analysis
[1967] 1. The device sends the video uploaded by the user to the server.
[1968] Users send video files and their meta information (title, description, etc.) to the server from their devices such as smartphones or PCs.
[1969] 2. The server receives the video, checks the quality, and generates metadata.
[1970] The server checks the format of the video file, converts it to a standard format if necessary, checks the video quality, and generates metadata (file size, duration, etc.).
[1971] 3. The server uses AI technology to analyze the video content and generate a summary.
[1972] The server analyzes the video frames to identify key scenes and important objects, and performs audio analysis to identify speakers and extract important phrases. Based on this information, a video summary is generated.
[1973] Categorizing the video
[1974] 1. The server categorizes the video based on the generated summary.
[1975] The server classifies videos into specific categories based on tags and keywords extracted from the summaries, and also assigns these tags and keywords to the video data.
[1976] 2. The server creates a recommendation list appropriate for each user.
[1977] The server selects the most suitable videos based on the user's viewing history and interests, and generates a recommendation list, which is then sent to the user's device.
[1978] Emotion analysis and recommendations using an emotion engine
[1979] 1. The server analyzes the user's emotions using an emotion engine.
[1980] The server analyzes the user's emotions based on facial expressions, tone of voice, viewing history, etc.
[1981] 2. The server reflects the results of the sentiment analysis in its video recommendations.
[1982] The server takes the emotional data into account and updates the recommendation list to provide more of the content the user enjoys.
[1983] Video Ad Optimization
[1984] 1. The server selects video ads based on user characteristics.
[1985] The server selects the most suitable advertisement based on the user's age, gender, interests, and emotional analysis results, and also sets the advertisement's meta information and playback timing.
[1986] 2. The server delivers ads within the video.
[1987] The server evaluates the relevance of the video content and inserts ads at the appropriate time, so that relevant ads are displayed when the user watches the video.
[1988] Determining what to monitor
[1989] 1. The server determines whether to monitor the video based on the summary.
[1990] The server detects whether the video summary contains inappropriate content (e.g., violence, adult content, etc.), performs a risk assessment, and determines whether manual supervision is necessary.
[1991] 2. The server will tag the data for manual monitoring as needed.
[1992] Videos that are deemed safe by automatic server monitoring will not be subject to manual monitoring. If manual monitoring is deemed necessary, the monitoring team will be notified.
[1993] Applying an age filter
[1994] 1. The server applies an age filter based on the user's estimated age.
[1995] The server estimates the user's age from their registration information and activity history, and detects age-restricted keywords (e.g., "violence" or "adult") from the video summary.
[1996] 2. The server filters inappropriate content.
[1997] The server prevents content that is not appropriate for the specified age from being played on the user's device, thereby preventing minors from accessing inappropriate content.
[1998] Specific examples
[1999] Example 1: User-Submitted Pet Videos
[2000] 1. The user (device) uploads a video of their dog playing at home.
[2001] 2. The server receives the video, performs quality checks, and generates metadata.
[2002] 3. The server analyzes the video and generates keywords such as "pet," "dog," and "play."
[2003] 4. The server uses these keywords to recommend relevant pet videos to the user.
[2004] 5. The server performs sentiment analysis on the user, detects whether the user is happy, and recommends more pet videos based on that.
[2005] Example 2: Fashion-related video ads
[2006] 1. The server identifies that the user has fashion-related interests.
[2007] 2. The server uses an emotion engine to confirm that the user has positive emotions toward the fashion content.
[2008] 3. The server selects fashion-related video ads and inserts them into the video at the appropriate time.
[2009] 4. The device displays relevant ads while you watch fashion videos.
[2010] Example 3: Reducing monitoring costs
[2011] 1. The server verifies the summary of the posted video for inappropriate content.
[2012] 2. If the server determines that there is no problem, the video will be automatically made public and will no longer be subject to manual monitoring.
[2013] Prompt Sentence Examples
[2014] 1. Analyze the video content and extract key scenes and important phrases.
[2015] 2. Generate a recommendation list based on the user's viewing history.
[2016] 3. Use a sentiment engine to analyze user sentiment and provide appropriate content.
[2017] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2018] Step 1:
[2019] The device sends the video file and its meta information (title, description, etc.) to the server.
[2020] Input: A user uploads a video file and...
Claims
1. A means by which videos are uploaded from the device, a means for checking the quality of the video received by the server and generating metadata; A means for the server to analyze the video content and generate a summary; The server categorizes the videos based on the generated summaries and assigns tags and keywords to them. A means for the server to add videos to a recommendation list based on user characteristics; A means for the server to select and provide the most suitable advertisement based on the user's characteristics; A means for the server to determine whether to monitor automatically or manually based on the video content; A means for the server to apply an age filter based on the estimated age of the User; A system including:
2. The system of claim 1 , wherein the terminal has a plurality of videos uploaded thereto.
3. The system according to claim 1, wherein the server uses AI technology for video analysis.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A