System

The system uses an AI engine to analyze and edit inappropriate video content, addressing parental challenges in managing children's video content, ensuring a safe viewing experience.

JP2026028702APending Publication Date: 2026-02-20SOFTBANK GROUP CORP

Patent Information

Application Number
JP2024131318
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

Parents face challenges in managing and restricting inappropriate content in videos watched by children, particularly violent scenes and inappropriate language, which can have a negative impact on children.

Method used

A system that utilizes an AI engine to analyze video frames and audio data, identify inappropriate content, tag such scenes, and allow parents to review and edit them automatically, converting inappropriate content into appropriate content using an editing module.

Benefits of technology

Provides a safe video environment for children by automatically identifying and editing inappropriate content, allowing parents to manage video content effortlessly and ensuring children watch suitable material.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026028702000001_ABST
    Figure 2026028702000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system, comprising: means for selecting or uploading a video that a user wants to view; a server for storing and analyzing video files; an artificial intelligence engine for extracting content of the analyzed video as feature vectors and identifying scenes containing inappropriate expressions; means for tagging the inappropriate scenes identified by the artificial intelligence engine and notifying a parent dashboard; means for a parent to confirm and approve modification of the inappropriate scenes; an editing module for converting the inappropriate scenes into appropriate content; and means for generating and providing an edited video to a user terminal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, anime and videos popular among children often contain content that is educationally inappropriate. It is extremely difficult for parents to review and restrict all of this content. In particular, violent scenes, inappropriate language, and depictions of bullying on social media can have a negative impact on children. The present invention aims to provide an environment where children can watch videos safely without parents having to manually review them, and to provide a means to solve this problem. [Means for solving the problem]

[0005] The present invention first provides a means for users (parents or children) to select or upload videos they wish to watch. The video files are then received by a server for storage and analysis. The server then requests an AI engine to analyze the video, extracting frame and audio data as feature vectors and using a specific algorithm to identify scenes containing inappropriate content.

[0006] The AI ​​engine then tags the inappropriate scenes it identifies and notifies the parent's dashboard. The parent can then review the content of the scenes and approve the edits. The video content is then instantly edited using an editing module to convert the inappropriate scenes into appropriate content. The edited video is then regenerated and provided to the parent's device. This allows parents to provide a safe video environment for their children to watch without any effort on their part.

[0007] "User" refers to a parent or child using the system who selects or uploads videos.

[0008] A "server" is a collection of computing resources that stores and analyzes video files, and generates and provides edited videos.

[0009] An "artificial intelligence engine" is software that implements algorithms that analyze video frames and audio data and identify scenes that contain inappropriate content.

[0010] The "dashboard" is an interface that allows parents to review the contents of a scene and approve modifications.

[0011] "Inappropriate scenes" are specific parts of videos that contain violence, inappropriate language, depictions of bullying, or other content that is deemed educationally inappropriate.

[0012] An "editing module" is a software component for converting inappropriate scenes into appropriate content.

[0013] An "analysis job" is a processing task for analyzing a video file using an artificial intelligence engine.

[0014] A "feature vector" is a collection of elements extracted from video frames or audio data, and is the information that is the subject of analysis. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] The present invention relates to a system for providing an environment in which users can safely and appropriately watch videos they wish to watch. As shown below, the processing contents of a program in which the server, terminal, and user each play an important role are described.

[0037] First, a user logs in to the system's web application and selects or uploads the video they want to watch. At this time, the device sends the selected video file to the server. The server stores the received video file in temporary storage and starts the analysis job.

[0038] The server then transfers the video file to an AI engine for analysis. The AI ​​engine analyzes the video frames and audio data to extract feature vectors. This determines whether the video contains inappropriate content. For example, scenes in which a character attacks another character with a weapon or scenes containing large amounts of blood are identified as inappropriate content.

[0039] The server then displays the timestamps and scene tags of inappropriate scenes identified by the AI ​​engine on the parent's dashboard. Parents can log in to the dashboard and review the inappropriate scenes. If the parent approves the modification of the inappropriate scene, the server launches an editing module to convert the inappropriate content into an appropriate one. For example, in a scene depicting a large amount of blood, the color of the blood can be changed to a harmless liquid, or the scene itself can be replaced with softer content.

[0040] Once the editing is complete, the video is generated as a new file, which the server provides to the parent's device. The parent can then play the edited video on the device and confirm that the child is safe to watch.

[0041] As a concrete example, consider the case where a parent uploads a "children's cartoon" to the system. The system receives the video and identifies violent scenes as a result of analysis. The violent scenes are displayed on the dashboard, and when the parent clicks the approval button, the scenes are modified. For example, in a scene where a character slashes an enemy with a sword, the sword is changed to a soft stick and the depiction of the attack itself is deleted. This edited video is then provided to the parent's device, allowing the child to watch it safely.

[0042] This invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. In addition, the automatic analysis and editing functions using AI make it possible to provide an appropriate video environment without any hassle.

[0043] The processing flow will be explained below.

[0044] Step 1:

[0045] Users log into the system's web application and select or upload the video they want to watch.

[0046] Step 2:

[0047] The terminal transmits the selected or uploaded video file to the server.

[0048] Step 3:

[0049] The server stores the received video file in temporary storage.

[0050] Step 4:

[0051] The server adds a job to the queue to start video analysis.

[0052] Step 5:

[0053] The server runs the analysis job and transfers the video file to the artificial intelligence engine.

[0054] Step 6:

[0055] The artificial intelligence engine analyzes video frames and extracts feature vectors for each frame.

[0056] Step 7:

[0057] The artificial intelligence engine also analyzes voice data and converts it into text data.

[0058] Step 8:

[0059] The artificial intelligence engine uses pre-trained models to determine whether or not there is inappropriate language.

[0060] Step 9:

[0061] The artificial intelligence engine then timestamps and tags identified inappropriate scenes.

[0062] Step 10:

[0063] The server notifies the parent's dashboard of the tagged inappropriate scenes.

[0064] Step 11:

[0065] The user (parent) logs in to the dashboard and checks for inappropriate scenes.

[0066] Step 12:

[0067] The user (parent) approves or rejects the correction of the displayed inappropriate scene.

[0068] Step 13:

[0069] The server, upon receiving parental approval, queues a job to convert the inappropriate scenes to appropriate content.

[0070] Step 14:

[0071] The server launches the editing module to process the scene to be modified.

[0072] For example, frames depicting large amounts of blood could be converted to harmless versions or the scene could be removed entirely.

[0073] Step 15:

[0074] The server also edits the audio data to the appropriate content, maintaining overall synchronization.

[0075] Step 16:

[0076] The server generates a new video file after the editing is complete.

[0077] Step 17:

[0078] The server uploads the edited video to the user's dashboard and notifies them.

[0079] Step 18:

[0080] The user (parent) can check the edited video from the dashboard and play it on their device.

[0081] Step 19:

[0082] The user (child) can safely watch the edited video.

[0083] Example 1

[0084] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0085] There are many videos on the Internet that contain inappropriate content for children, making it difficult for parents to manage the content their children view. Editing out inappropriate scenes individually takes time and effort. To solve this problem, a system is needed that can automatically detect inappropriate content and edit it safely.

[0086] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0087] In this invention, the server includes a means for users to select or upload videos they wish to watch, a means for saving the video files in temporary storage and creating analysis jobs, and a means for transferring the video files to an artificial intelligence engine and requesting analysis. This makes it possible to automatically analyze videos containing inappropriate content and edit them in a safe manner.

[0088] "User" means an individual or entity that uses the System to select or upload videos that they wish to view.

[0089] A "video file" is a digital data file that contains visual and audio content.

[0090] "Temporary storage" is a digital storage device for temporarily storing video files until an analysis job is started.

[0091] An "analysis job" is a processing task that transfers a video file to an artificial intelligence engine and has its contents analyzed.

[0092] An "artificial intelligence engine" is a software or hardware system that analyzes video frames and audio data and runs algorithms to identify inappropriate language.

[0093] A "feature vector" is data that expresses each feature extracted from video frames or audio as a numerical value.

[0094] "Inappropriate content" refers to video scenes that contain violence, blood, or obscene content that is undesirable for children or general audiences.

[0095] A "timestamp" is data indicating the occurrence time of a frame containing an inappropriate expression.

[0096] A "scene tag" is a label that is assigned to identify information related to a particular scene.

[0097] "Parents" are guardians who are responsible for managing the content their children view.

[0098] The "overview panel" is an interface provided by the system, which is a dashboard for users to check the video analysis results and inappropriate scenes.

[0099] An "editing module" is a software or hardware system for changing detected inappropriate scenes into appropriate content.

[0100] "User Terminal" means a device used by a parent or user to access the system and view or manage videos.

[0101] A "new file" is a digital file generated containing the video modified by the editing module.

[0102] The present invention provides a system for providing an environment in which users can safely and appropriately watch videos they wish to watch. This system is realized using the following specific hardware and software.

[0103] 1. User video selection and upload

[0104] Users log in to the system's web application and select or upload the video they want to watch. To upload, the device sends the video file to the server using an HTTP POST request.

[0105] 2. Video transmission and storage to the server

[0106] The video file sent by the device is received by the server and saved in temporary storage (e.g., AWS S3 bucket, Google Cloud Storage), ensuring safe storage until the analysis job begins.

[0107] 3. Start video analysis

[0108] The server obtains the path of the saved video file, creates an analysis job, and adds it to the queue. A message broker such as RabbitMQ or Apache Kafka is used for the analysis job queue. The server then transfers the analysis job to the AI ​​engine and requests it to be analyzed.

[0109] 4. Video analysis using an AI engine

[0110] The AI ​​engine splits the video file into frames and analyzes the frames and audio data in parallel using deep learning models (e.g., YOLO, ResNet, and WaveNet for audio analysis). It extracts feature vectors from each frame and audio and determines whether the features represent inappropriate content (e.g., violence, blood, obscenity).

[0111] 5. Dashboard display of inappropriate scene information

[0112] The server generates data to display the analysis results from the AI ​​engine on the parent's dashboard. When the parent logs in, this data is displayed and the parent can check the timestamps of inappropriate scenes and scene tags.

[0113] 6. Parental approval of inappropriate scenes and corrections

[0114] Parents can check inappropriate scenes via the dashboard and approve corrections by sending a correction request to the server.

[0115] 7. Edit video with the editing module

[0116] After receiving parental approval, the server launches the editing module, which applies filters and replacements to inappropriate scenes based on the AI ​​engine's judgment. These modifications include changing the color of blood and muting obscene content.

[0117] 8. Video generation and delivery to devices

[0118] The modified video file is created as a new file and saved in temporary storage. The server sends a link to this file to the parent device and notifies them that editing is complete.

[0119] Specific examples

[0120] When uploading a "children's cartoon," the system analyzes the video and identifies violent scenes. If the parent reviews the details of the violent scene displayed on the dashboard and approves the edits, a scene in which a character slashes an enemy with a sword is changed to a soft stick, and the attack is deleted. A new file is generated and provided to the parent's device. An example prompt is, "Please explain the system for editing and providing children's videos to make them safe."

[0121] This invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. In addition, the automatic analysis and editing functions using AI make it possible to provide an appropriate video environment without any hassle.

[0122] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0123] Step 1:

[0124] User selects or uploads a video

[0125] A user logs in to a web application and selects or uploads a video they want to watch. When the user clicks the upload button for a video file, the device sends the video file to the server using an HTTP POST request. The input is the video file selected by the user, and the output is the transmission of the video file to the server. Specifically, the device reads the metadata of the video file, generates an upload request, and sends it to the server.

[0126] Step 2:

[0127] Sending and saving video files to the server

[0128] The video file sent from the device is received by the server. The server saves the received video file in temporary storage (e.g., AWS S3 bucket, Google Cloud Storage). The input is the video file sent from the device, and the output is the video file saved in temporary storage. Specifically, the server generates a path to save the video file and uploads the file to storage.

[0129] Step 3:

[0130] Starting a video analysis job

[0131] The server obtains the path of the saved video file, creates an analysis job, and adds it to a message queue. The input is the path of the video file saved in temporary storage, and the output is the job added to the analysis job queue. Specifically, the server generates a job ID and adds the job details, along with the path of the video file, to a queue (such as RabbitMQ or Apache Kafka).

[0132] Step 4:

[0133] Video analysis by AI engine

[0134] The server transfers the analysis job to the AI ​​engine and requests analysis. The AI ​​engine divides the video file into frames and analyzes the frame and audio data. The input is the analysis job details (video file path, job ID), and the output is the analysis result data (feature vector, timestamp, scene tag). Specifically, the AI ​​engine analyzes the images and audio using YOLO, ResNet, or WaveNet models, and extracts feature vectors from each frame and audio.

[0135] Step 5:

[0136] Dashboard display of inappropriate scene information

[0137] The server converts the analysis results received from the AI ​​engine into dashboard data and displays it on the parent dashboard. The input is the analysis result data of inappropriate scenes, and the output is the timestamp and scene tag displayed on the parent dashboard. Specifically, the server formats the data and sends it to the dashboard frontend using WebSocket or HTTP requests.

[0138] Step 6:

[0139] Parental review and approval of inappropriate scenes

[0140] Parents log in to the dashboard, check the inappropriate scenes, and approve their corrections. The input is a list of inappropriate scenes displayed on the dashboard, and the output is a request for approval of the corrections. Specifically, when a parent checks the scene details and thumbnails and clicks the "Approve Corrections" button, a request is sent to the server.

[0141] Step 7:

[0142] Editing module for video editing

[0143] The server launches the editing module after receiving approval for the correction from the parent. The editing module corrects the inappropriate scenes based on the AI ​​engine's judgment. The input is the correction approval request and details of the inappropriate scenes, and the output is the corrected video file. Specifically, the editing module detects inappropriate frames and applies filters and replacement processes to correct the content of the scenes.

[0144] Step 8:

[0145] Generate edited video and deliver it to the device

[0146] The server generates a new video file modified by the editing module and stores it in temporary storage. The input is the modified video data, and the output is a link to the new video file. Specifically, the server generates the video file and sends its path to the parent device.

[0147] Step 9:

[0148] Play and check the edited video

[0149] The parent uses the provided link to play the edited video on their device and confirm the content. The input is a link to the edited video file, and the output is the confirmed video playback. Specifically, the parent clicks the link to launch the video player and watches the edited video to confirm that it is safe for their child to watch.

[0150] (Application example 1)

[0151] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0152] With the recent spread of video streaming services, the safety of children's content has become a major issue. In particular, it is important to prevent videos containing inappropriate content from falling into the hands of children. However, it takes a great deal of time and effort for parents to manually review and edit the content of all videos. To solve this problem, an efficient and highly accurate automated system is needed.

[0153] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0154] In this invention, the server includes: means for a user to select or upload a video they wish to watch; a server for storing and analyzing video files; an artificial intelligence engine for extracting the analyzed video content as a feature vector and identifying scenes containing inappropriate content; means for tagging the inappropriate scenes identified by the artificial intelligence engine and notifying the user on their dashboard; means for the user to review the inappropriate scenes and approve their correction; an editing module for converting the inappropriate scenes into appropriate content; and means for generating edited videos and providing them to a content distribution service. This allows users to quickly and easily manage, edit, and provide safe videos for children without any hassle.

[0155] "Means for users to select or upload videos they wish to view" refers to a function that allows users to use their own devices to specify or upload videos they wish to view to the system.

[0156] A "server for storing and analyzing video files" is a computer device that temporarily stores uploaded video files and analyzes their contents using an artificial intelligence engine.

[0157] The "artificial intelligence engine for extracting the analyzed video content as feature vectors and identifying scenes containing inappropriate expressions" is a processing device equipped with an artificial intelligence algorithm for analyzing video frames and audio data and automatically detecting inappropriate scenes based on the resulting feature vectors.

[0158] "Means for tagging inappropriate scenes identified by the AI ​​engine and notifying the user on their dashboard" refers to a mechanism for tagging inappropriate scenes detected by the AI ​​engine and displaying that information on a dashboard where users can view it.

[0159] "Means for users to confirm inappropriate scenes and approve corrections" refers to a function that allows users to use the dashboard to confirm the content of inappropriate scenes and select and approve whether or not to correct the scenes.

[0160] The "editing module for converting inappropriate scenes into appropriate content" is a software module that has an editing function for converting scenes that are determined to be inappropriate into more appropriate content based on the user's approval.

[0161] "Means for generating edited video and providing it to a content distribution service" refers to a function for generating a new video file containing the modified scene and providing that video to users via a content distribution service.

[0162] The present invention relates to a system that provides a safe viewing environment for video content for children. In this system, the server, the terminal, and the user each play an important role.

[0163] First, a user uploads or selects a video they wish to watch using their device. The video file is then sent to the server and temporarily stored in storage. The server then transfers the received video file to an AI engine for analysis. The AI ​​engine analyzes the video frames and audio data to extract feature vectors. This allows it to identify scenes that contain inappropriate content, such as violent scenes or scenes containing inappropriate language.

[0164] The analysis results are extracted as timestamps and scene tags for inappropriate scenes identified by the AI ​​engine. This information is then posted to the user's dashboard. Parents or appropriate administrators can log in to the dashboard to review the inappropriate scenes and approve their corrections. For example, if a character uses a sword in a violent scene, the scene can be changed to a soft stick, or the scene itself can be replaced with different content.

[0165] Once approval is given, an editing module is launched on the server side and the inappropriate scenes are converted into appropriate content. This editing process is carried out automatically using artificial intelligence. The edited video is generated as a new file and provided to the user's device via the content distribution service. This ensures that children can watch videos safely.

[0166] The following hardware and software can be used to build the system. It is recommended that the server be equipped with storage, a CPU, and a GPU for storing and analyzing video files. Deep learning frameworks such as TensorFlow and PyTorch can be used as the artificial intelligence engine. Terminals include devices such as smartphones, tablets, and PCs.

[0167] As a concrete example, consider the case where a user uploads a children's cartoon. The system receives the video and identifies violent scenes as a result of analysis by an artificial intelligence engine. The dashboard displays timestamps and scene tags for violent scenes, and parents can click a correction button to edit the scene. For example, in a scene where a character slashes an enemy with a sword, the sword can be changed to a soft stick. In this way, an edited video is generated and provided to the user's device via a distribution service.

[0168] Examples of prompts include:

[0169] In order to analyze videos for inappropriate content and provide a safe viewing environment, please analyze and edit your videos according to the following requirements.

[0170] Identify any inappropriate scenes and provide their start and end times

[0171] If a scene needs editing, replace it with appropriate content.

[0172] In this way, users can quickly and easily manage, edit and provide safe videos for children without any hassle.

[0173] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0174] Step 1:

[0175] A user uses a terminal to upload the video they want to watch to the system. The input is the video file selected or uploaded by the user, and the output is the video file sent to the server, at which point the file is saved in temporary storage.

[0176] Step 2:

[0177] The server transfers the received video file to the AI ​​engine and begins analysis. The input is the video file stored on the server, and the output is the video data sent to the AI ​​engine. The server completes the video file transfer procedure and prepares for analysis.

[0178] Step 3:

[0179] The AI ​​engine analyzes video frames and audio data to extract feature vectors. The input is the video data transferred from the server, and the output is the analysis result: feature vectors, timestamps of inappropriate scenes, and scene tags. An AI model (e.g., TensorFlow or PyTorch) is used to analyze each frame in the video and perform data calculations to detect inappropriate content.

[0180] Step 4:

[0181] The server tags inappropriate scenes identified by the AI ​​engine and notifies the user's dashboard. The input is the analysis result from the AI ​​engine, and the output is the inappropriate scene information notified on the dashboard. The server reflects the tag information and timestamp on the dashboard.

[0182] Step 5:

[0183] The user logs in to the dashboard, checks the inappropriate scenes, and approves the corrections. The input is the inappropriate scene information displayed on the dashboard, and the output is a command to approve the corrections. The user's operation determines whether the inappropriate scenes need to be corrected.

[0184] Step 6:

[0185] The server starts the editing module based on the approved correction request. The input is a command from the user to approve the correction, and the output is the start of the editing module. The editing module automatically converts inappropriate scenes into appropriate ones.

[0186] Step 7:

[0187] The editing module converts inappropriate scenes into appropriate content and generates a new edited video file. The input is the inappropriate scene data handled by the editing module, and the output is the edited video file. The AI ​​model is used to correct the content of the scene and generate a new edited video.

[0188] Step 8:

[0189] The server saves the edited video file as a final file and provides it to the user's device through the content distribution service. The input is the edited video file, and the output is the video file provided to the user's device through the distribution service. Users can receive safe video content for children to watch.

[0190] This allows users to easily create a safe video viewing environment for children.

[0191] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0192] The present invention relates to a system for providing an environment in which users can watch videos they want to watch in a safe and appropriate manner. In particular, the present invention describes a system that combines an emotion engine that recognizes the user's emotions and efficiently identifies and corrects inappropriate scenes.

[0193] The system begins with the user selecting or uploading a video they want to watch. Once the user uploads the video, the device sends the video file to the server. The server then stores the received video file in temporary storage and starts the analysis job.

[0194] The server transfers the video file to an AI engine, which extracts feature vectors from the video frames and audio data. This allows a pre-trained model to determine whether the video contains inappropriate content. For example, scenes in which a character attacks another character with a weapon or contains a large amount of blood are identified as inappropriate content.

[0195] Furthermore, the present invention incorporates an emotion engine that collects emotional data from the user while watching. The emotion engine acquires emotional data such as the user's facial expressions and heart rate through cameras and sensors, and associates the data with inappropriate scenes. For example, if the user's heart rate temporarily rises or if a stressed reaction is observed in the user's facial expression, the emotion engine determines that the scene is likely to be inappropriate.

[0196] The server displays the timestamps and scene tags of inappropriate scenes identified by the AI ​​engine, as well as emotional data from the emotion engine, on the parent's dashboard. Parents log in to the dashboard and review the inappropriate scenes. If the parent approves the correction of a confirmed scene, the server launches an editing module to convert the inappropriate content into something appropriate. For example, in a scene depicting a large amount of blood, the color of the blood can be changed to a harmless liquid, or the scene itself can be replaced with softer content.

[0197] The edited video is then generated as a new file, which the server then provides to the parent's device. The parent then plays the edited video to ensure it is safe for their child to watch. The emotion engine also collects emotional data when the parent approves or rejects the edits, allowing the AI ​​engine to learn from this data to help identify inappropriate scenes in the future.

[0198] As a concrete example, consider the case where a parent uploads a "family movie" to the system. The system receives the video and analyzes it using AI and an emotion engine. The emotion engine collects user emotional data and determines violent or tragic scenes as inappropriate. Based on this information, the parent approves the scene modification on the dashboard, and the AI ​​engine converts the scene into appropriate content, resulting in a safe video.

[0199] In this way, this invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. Furthermore, by combining AI and an emotion engine, it is possible to identify inappropriate scenes with greater accuracy and make appropriate corrections.

[0200] The processing flow will be explained below.

[0201] Step 1:

[0202] Users log into the system's web application and select or upload the video they want to watch.

[0203] Step 2:

[0204] The terminal transmits the selected or uploaded video file to the server.

[0205] Step 3:

[0206] The server stores the received video file in temporary storage.

[0207] Step 4:

[0208] The server adds a job to the queue to start video analysis.

[0209] Step 5:

[0210] The server runs the analysis job and transfers the video file to the artificial intelligence engine.

[0211] Step 6:

[0212] The artificial intelligence engine analyzes video frames and extracts feature vectors for each frame.

[0213] Step 7:

[0214] The artificial intelligence engine analyzes the voice data and converts it into text data.

[0215] Step 8:

[0216] The artificial intelligence engine uses pre-trained models to determine whether a sentence contains inappropriate language.

[0217] Step 9:

[0218] The artificial intelligence engine will timestamp and tag inappropriate scenes.

[0219] Step 10:

[0220] The emotion engine collects emotional data while users watch videos, for example by analyzing their facial expressions using a camera and monitoring their heart rate with a heart rate sensor.

[0221] Step 11:

[0222] The emotion engine analyzes the collected emotion data and associates it with inappropriate scenes.

[0223] Step 12:

[0224] The server notifies the parent's dashboard of tagged inappropriate scenes and associated emotional data.

[0225] Step 13:

[0226] The user (parent) logs in to the dashboard and checks the inappropriate scenes and their emotion data.

[0227] Step 14:

[0228] The user (parent) approves or rejects the correction of the inappropriate scene.

[0229] Step 15:

[0230] The server, upon receiving parental approval, queues a job to convert the inappropriate scenes to appropriate content.

[0231] Step 16:

[0232] The server launches an editing module to process the scene to be altered, for example, converting a frame depicting large amounts of blood into a harmless liquid and replacing the scene with a different musical melody.

[0233] Step 17:

[0234] The server also edits the audio data to the appropriate content, maintaining overall synchronization.

[0235] Step 18:

[0236] The server generates a new video file after the editing is complete.

[0237] Step 19:

[0238] The server uploads the edited video to the user's dashboard and notifies them.

[0239] Step 20:

[0240] The user (parent) can check the edited video from the dashboard and play it on their device.

[0241] Step 21:

[0242] The user (child) can safely watch the edited video.

[0243] Step 22:

[0244] The emotion engine also collects emotional data when parents approve or reject modifications and feeds this back to the AI ​​engine.

[0245] Step 23:

[0246] The artificial intelligence engine uses the collected emotional data to optimize algorithms for identifying inappropriate scenes.

[0247] Example 2

[0248] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0249] Currently, when a scene containing inappropriate content is found in video content viewed by children, parents must manually detect and edit that scene. This places a significant burden on parents, and there are limitations to the accuracy of identifying and editing inappropriate scenes. Furthermore, there is no established method for effectively utilizing user emotion data to improve the accuracy of identifying inappropriate scenes. Therefore, there is a need for a system that provides a safe and appropriate video viewing environment and reduces the burden on parents.

[0250] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for selecting or uploading a video that a user wants to watch; means for saving and analyzing video files; means for extracting the content of the analyzed video as a feature vector and identifying scenes containing inappropriate expressions; means for collecting user emotion data and associating it with inappropriate scenes; means for tagging inappropriate scenes identified by the artificial intelligence engine and the emotion recognition engine and notifying them to a parent's dashboard; means for a parent to confirm and approve corrections of inappropriate scenes; means for converting inappropriate scenes into appropriate content; and means for generating edited videos and providing them to a user terminal. This enables automatic detection and editing of scenes containing inappropriate expressions, significantly reducing the burden on parents and providing an environment in which children can watch safe and appropriate videos.

[0251] "User" refers to the person who selects or uploads videos to the system that they wish to view.

[0252] "Video File" means a digital file containing video and audio data selected or uploaded by a User for viewing purposes.

[0253] A "server" refers to a computer system that stores video files and handles a series of processes such as analyzing and editing them.

[0254] "Feature vector" refers to vector data that quantifies specific attributes and features extracted from video frames and audio.

[0255] "Artificial intelligence engine" refers to a machine learning model that uses feature vectors to analyze videos and identify scenes containing inappropriate content.

[0256] An "emotion recognition engine" refers to a system that uses cameras and sensors to collect user emotional data and associates it with inappropriate scenes.

[0257] "Dashboard" refers to the user interface that allows parents to access the system to review information about inappropriate scenes and approve corrections.

[0258] "Editing Module" refers to software functionality for converting identified inappropriate scenes into appropriate content.

[0259] A "timestamp" refers to data that indicates the specific time at which a particular frame or scene appears in a video.

[0260] "Tags" refer to metadata used to classify and identify inappropriate scenes.

[0261] The present invention relates to a system for providing an environment in which a user can safely and appropriately view videos that the user desires to view. Detailed embodiments of the present invention will be described below.

[0262] First, the user selects or uploads the video they want to watch. The device sends the video file uploaded by the user to the server. The device can be a general personal computer or smartphone, and uses a file selection dialog that runs on a web browser.

[0263] Next, the server temporarily stores the received video file in a cloud storage service such as Amazon S3.

[0264] The server queues an analysis job for the stored video file and starts executing the job. The server sends the video file to an artificial intelligence engine (e.g., a TensorFlow model) that extracts feature vectors from the video frames and audio data.

[0265] The AI ​​engine uses feature vectors, specifically pre-trained machine learning models, to identify inappropriate scenes, such as scenes of violence or excessive gore.

[0266] Furthermore, the system collects user emotional data using cameras and sensors connected to the device. A standard webcam (e.g., Logitech C920) is used as the camera, and a heart rate monitor (e.g., Fitbit Charge 4) is used as the sensor. The Affectiva SDK is used as the emotion recognition engine. This allows the system to capture the user's facial expressions and heart rate in real time and associate them with inappropriate scenes.

[0267] The server displays the collected emotion data and timestamps and tags of inappropriate scenes on a dashboard. The dashboard is implemented using a web framework such as React.js. Parents can log in to the dashboard and check the inappropriate scenes.

[0268] If a parent approves the correction of an inappropriate scene, the server launches an editing module (e.g., FFmpeg command) that converts the inappropriate content into something appropriate, such as changing the color of blood or cutting the scene altogether.

[0269] Once the editing is complete, the video is generated as a new file, which is then provided to the device by the server. A link is displayed on the dashboard for parents to download the new video file. When parents click the link, the file is downloaded to the device.

[0270] Finally, parents can play the edited video on their device to confirm that it is safe for their child to watch. The emotion recognition engine also collects emotional data when parents approve or reject the edits, and uses this data for future learning on the server. This improves the accuracy of identifying inappropriate scenes in future videos.

[0271] As a concrete example, consider the case where a parent uploads a "family movie" to the system. The system receives the video and analyzes it using AI and an emotion recognition engine. For example, it uses a prompt such as, "Please detect whether this movie contains violent or tragic scenes." The emotion recognition engine collects the user's emotional data and determines that violent or tragic scenes are inappropriate. Based on this information, if the parent approves the scene modification on the dashboard, the AI ​​engine converts the scene into appropriate content, and a safe video is generated.

[0272] In this way, this invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. Furthermore, by combining AI and an emotion recognition engine, it is possible to identify inappropriate scenes with greater accuracy and make appropriate corrections.

[0273] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0274] Step 1:

[0275] The user selects or uploads the video they want to watch. The device retrieves the video file selected or uploaded by the user and sends it to the server using an HTTP POST request. The input is the video file selected by the user, and the output is the video file sent to the server.

[0276] Step 2:

[0277] The server saves the received video file in temporary storage. Specifically, it uploads the video file to cloud storage such as Amazon S3 and records the status of the saving completion in a log. The input is the video file sent to the server, and the output is the video file saved in the storage.

[0278] Step 3:

[0279] The server queues an analysis job for the stored video file and starts the job execution. The video file is sent to the artificial intelligence engine for analysis. The input is the video file stored in storage, and the output is the start status of the analysis job.

[0280] Step 4:

[0281] The server transfers the video file to an AI engine, which extracts feature vectors from the video frames and audio data. The AI ​​engine (e.g., a TensorFlow model) converts each video frame and audio data into a numerical vector and inputs it into a generative AI model. The input is the video file and its frames and audio data, and the output is the feature vector.

[0282] Step 5:

[0283] The AI ​​engine uses feature vectors to identify inappropriate scenes. Specifically, it applies a pre-trained model and uses a specific algorithm to detect frames containing inappropriate content. The input is the feature vectors, and the output is a list of inappropriate scenes with timestamps and tags.

[0284] Step 6:

[0285] The device uses an emotion recognition engine to collect user emotional data. It uses cameras and sensors connected to the device to capture the user's facial expressions and heart rate in real time. The collected data is sent to a server. The input is raw data from the video and heart rate sensors, and the output is emotional data collected in real time.

[0286] Step 7:

[0287] The server analyzes the collected emotion data and associates it with inappropriate scenes. The emotion recognition engine analyzes the collected data and identifies emotional changes in specific scenes. For example, it identifies when a person's heart rate spikes and associates this with inappropriate scenes. The input is emotion data, and the output is emotion data associated with inappropriate scenes.

[0288] Step 8:

[0289] The server displays the timestamps and tags of inappropriate scenes, as well as emotion data, on a dashboard. It uses React.js to display the information in a web interface that parents can access. The input is inappropriate scenes and emotion data, and the output is the data displayed on the dashboard.

[0290] Step 9:

[0291] The parent logs in to the dashboard and checks the inappropriate scenes. After checking, the parent "approves" or "rejects" the scene correction. The input is the data displayed on the dashboard, and the output is the parent's operation data.

[0292] Step 10:

[0293] The server starts the editing module and converts the content of the inappropriate scene into an appropriate one. Specifically, it uses FFmpeg commands to process the scene that needs to be edited. The input is the parent operation data and the inappropriate scene information, and the output is the edited scene.

[0294] Step 11:

[0295] The server generates a new video file containing the modified content. A new edited video file is created and saved to storage. The input is the edited scene, and the output is the new video file.

[0296] Step 12:

[0297] The server provides the generated new video file to the parent's device. A download link is generated on the dashboard, which the parent can access to download the new video. The input is the new video file, and the output is the download link.

[0298] Step 13:

[0299] The parent plays the edited video on the device and confirms that it is safe for the child to watch. The input is the new video file, and the output is the parent's confirmation result.

[0300] Step 14:

[0301] The emotion recognition engine also collects emotional data when parents approve or reject corrections, and the server uses this data for future learning. This improves the accuracy of identifying inappropriate scenes from the next time onwards. The input is the parent's emotional data, and the output is learning data.

[0302] (Application example 2)

[0303] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0304] In conventional digital content viewing systems, it is difficult for users to manually detect and correct inappropriate scenes in the content they are viewing. Simply deleting inappropriate scenes often detracts from the user's viewing experience. Furthermore, there is a lack of a mechanism for more precisely identifying inappropriate scenes using user emotional data. This makes it difficult for users to enjoy content safely and comfortably.

[0305] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0306] In this invention, the server includes a means for a user to select or upload digital content they wish to view, a means for saving and analyzing digital content files, and a machine learning engine that extracts the analyzed digital content content as a feature vector and identifies scenes containing inappropriate content. This makes it possible to identify inappropriate scenes using emotion data. The server also includes a means for tagging inappropriate scenes identified by the machine learning engine and notifying a parental interface, a means for a parent to confirm the inappropriate scenes and approve corrections, and an editing module for converting the inappropriate scenes into appropriate content, thereby providing a safe and comfortable viewing experience for users. Furthermore, the server includes a means for collecting user emotion data and associating it with inappropriate scenes using the emotion engine, and a means for collecting emotion data when a parent approves or rejects corrections and performing learning based on this data, making it possible to more precisely identify inappropriate scenes and make appropriate corrections.

[0307] "Digital content" means media, including video, audio, images, and text, that is stored and distributed in electronic form.

[0308] A "machine learning engine" is a program that contains algorithms that analyze data, automatically learn patterns and features without relying on human-defined rules, and make predictions and classifications.

[0309] "Parental interface" refers to a user interface used by a parent to notify the parent of an inappropriate scene, and to review and approve the correction.

[0310] "Inappropriate content" refers to scenes that contain content that may have an upsetting or harmful effect on viewers, such as violence, discrimination, or extreme tragedy.

[0311] The "emotion engine" is a system that uses cameras and sensors to acquire biometric information such as a user's facial expressions and heart rate, and analyzes their emotional state.

[0312] An "editing module" is a software component for converting the content of identified inappropriate scenes into appropriate content.

[0313] A "feature vector" is data that quantifies the content extracted from the frames and audio data of digital content.

[0314] An "algorithm" is a set of steps or a computational method for solving a particular problem.

[0315] A "timestamp" is data that indicates the time at which a particular frame or scene occurred.

[0316] "Learning means" is the process by which a machine learning engine improves its model based on collected data.

[0317] MODE FOR CARRYING OUT THE INVENTION

[0318] To implement this invention, the following configuration and operation procedures are required.

[0319] The system of the present invention includes a terminal for users to select or upload digital content they wish to view, and a server for storing and analyzing digital content files. The server is equipped with a machine learning engine and an emotion engine, and collects and analyzes user emotion data in real time.

[0320] First, a user uploads digital content to a device. The device then sends the digital content file to a server. The server stores the received digital content file in temporary storage and initiates an analysis job. The server then transfers the digital content file to a machine learning engine, which extracts feature vectors from video frames and audio data. This allows a pre-trained algorithm to determine whether the content contains inappropriate language.

[0321] At the same time, the user's device is equipped with a camera and sensors, which allow the emotion engine to capture emotional data such as the user's facial expressions and heart rate. When an inappropriate scene is identified, the scene's timestamp and scene tag are recorded. When an inappropriate scene is identified, this information is notified to the parental interface.

[0322] Parents can review inappropriate scenes and approve their edits via the interface. Once the edits are approved, the server launches an editing module to convert the inappropriate content into appropriate content. For example, in a scene that contains a large amount of blood, the color of the blood can be changed to a harmless liquid, or the scene itself can be replaced with softer content.

[0323] Once edits are complete, the digital content is generated as a new file, which the server then provides to the user's device. The user can then play the edited digital content for safe and comfortable viewing. The emotion engine also collects emotional data from parents when they approve or reject the edits, which the machine learning engine can use to learn from to help identify inappropriate scenes in the future.

[0324] As a concrete example, consider the case where a children's animated content is uploaded. The system analyzes every frame and detects violent or tragic scenes. Parents can then review the content of the scenes on the dashboard and approve the edits. The edited video can then be safely shown to children.

[0325] An example of a prompt statement can be specified as follows:

[0326] Given the following text, generate a step-by-step guide to detecting and correcting inappropriate scenes in a children's cartoon, including collecting emotion data and analyzing the scene using AI.

[0327] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0328] Step 1:

[0329] The user selects or uploads digital content on the device.

[0330] Input: Digital content files (video files, audio files, etc.)

[0331] Output: The digital content file is saved to the server.

[0332] Specific operation: The user selects digital content through the terminal interface and clicks the upload button, after which the file is sent to the server.

[0333] Step 2:

[0334] The server stores the digital content file in temporary storage and starts the analysis job.

[0335] Input: Digital content file submitted by the user

[0336] Output: The digital content file is saved to temporary storage and analysis begins.

[0337] What happens: The server receives the digital content file and automatically stores it in temporary storage, then prepares it for analysis script execution.

[0338] Step 3:

[0339] The server transfers the digital content files to a machine learning engine, which extracts feature vectors for the video frames and audio data.

[0340] Input: Digital content file

[0341] Output: Feature vector (frame data, numerical representation of audio data)

[0342] Specific operation: The server transfers the digital content file to the machine learning engine, which processes the file to extract feature vectors from the frame data and audio data.

[0343] Step 4:

[0344] A machine learning engine uses feature vectors to detect scenes containing inappropriate language and record them with a timestamp.

[0345] Input: feature vector

[0346] Output: timestamp and tag data for the incorrect scene

[0347] How it works: The machine learning engine uses trained algorithms to analyze feature vectors and identify scenes that contain inappropriate content, such as violence or discrimination.

[0348] Step 5:

[0349] The device collects user emotional data in real time through cameras and sensors, and the emotion engine analyzes it.

[0350] Input: User's facial expression data, heart rate data

[0351] Output: Emotional data (stress response, heart rate variability, etc.)

[0352] Specific operation: The device's camera and sensors collect the user's facial expressions and biometric information, which are then sent to the emotion engine for analysis.

[0353] Step 6:

[0354] The emotion engine associates it with an inappropriate scene and sends it to the server.

[0355] Input: Emotion data, timestamps and tag data of inappropriate scenes

[0356] Output: Emotion data and association information for inappropriate scenes

[0357] Specific operation: The emotion engine matches the user's emotion data with the timestamp of the inappropriate scene based on the log.

[0358] Step 7:

[0359] The parental interface will be notified of inappropriate scenes and asked for approval to make corrections.

[0360] Input: timestamps and tag data of inappropriate scenes, emotion data

[0361] Output: Parental notification, request for approval of corrections

[0362] Specific behavior: The server displays detailed information about the inappropriate scene in the parental interface and asks the parent to approve or reject the correction.

[0363] Step 8:

[0364] After receiving parental approval, the server launches an editing module and converts the inappropriate scenes into appropriate content.

[0365] Input: Parental edit approval, timestamp and tag data for inappropriate scenes

[0366] Output: Edited scene data

[0367] Specific operation: The server's editing module converts the specified inappropriate scene into an appropriate expression and generates new scene data.

[0368] Step 9:

[0369] The edited digital content is generated as a new file and provided to the user's terminal.

[0370] Input: Edited scene data, original digital content excluding inappropriate scenes

[0371] Output: Edited digital content file

[0372] Specific operation: The server integrates the edited scene data into the entire original content, generates a new digital content file, and sends it to the user's device.

[0373] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0374] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0375] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0376] [Second embodiment]

[0377] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0378] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0379] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0380] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0381] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0382] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0383] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0384] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0385] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0386] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0387] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0388] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0389] The present invention relates to a system for providing an environment in which users can safely and appropriately watch videos they wish to watch. As shown below, the processing contents of a program in which the server, terminal, and user each play an important role are described.

[0390] First, a user logs in to the system's web application and selects or uploads the video they want to watch. At this time, the device sends the selected video file to the server. The server stores the received video file in temporary storage and starts the analysis job.

[0391] The server then transfers the video file to an AI engine for analysis. The AI ​​engine analyzes the video frames and audio data to extract feature vectors. This determines whether the video contains inappropriate content. For example, scenes in which a character attacks another character with a weapon or scenes containing large amounts of blood are identified as inappropriate content.

[0392] The server then displays the timestamps and scene tags of inappropriate scenes identified by the AI ​​engine on the parent's dashboard. Parents can log in to the dashboard and review the inappropriate scenes. If the parent approves the modification of the inappropriate scene, the server launches an editing module to convert the inappropriate content into an appropriate one. For example, in a scene depicting a large amount of blood, the color of the blood can be changed to a harmless liquid, or the scene itself can be replaced with softer content.

[0393] Once the editing is complete, the video is generated as a new file, which the server provides to the parent's device. The parent can then play the edited video on the device and confirm that the child is safe to watch.

[0394] As a concrete example, consider the case where a parent uploads a "children's cartoon" to the system. The system receives the video and identifies violent scenes as a result of analysis. The violent scenes are displayed on the dashboard, and when the parent clicks the approval button, the scenes are modified. For example, in a scene where a character slashes an enemy with a sword, the sword is changed to a soft stick and the depiction of the attack itself is deleted. This edited video is then provided to the parent's device, allowing the child to watch it safely.

[0395] This invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. In addition, the automatic analysis and editing functions using AI make it possible to provide an appropriate video environment without any hassle.

[0396] The processing flow will be explained below.

[0397] Step 1:

[0398] Users log into the system's web application and select or upload the video they want to watch.

[0399] Step 2:

[0400] The terminal transmits the selected or uploaded video file to the server.

[0401] Step 3:

[0402] The server stores the received video file in temporary storage.

[0403] Step 4:

[0404] The server adds a job to the queue to start video analysis.

[0405] Step 5:

[0406] The server runs the analysis job and transfers the video file to the artificial intelligence engine.

[0407] Step 6:

[0408] The artificial intelligence engine analyzes video frames and extracts feature vectors for each frame.

[0409] Step 7:

[0410] The artificial intelligence engine also analyzes voice data and converts it into text data.

[0411] Step 8:

[0412] The artificial intelligence engine uses pre-trained models to determine whether or not there is inappropriate language.

[0413] Step 9:

[0414] The artificial intelligence engine then timestamps and tags identified inappropriate scenes.

[0415] Step 10:

[0416] The server notifies the parent's dashboard of the tagged inappropriate scenes.

[0417] Step 11:

[0418] The user (parent) logs in to the dashboard and checks for inappropriate scenes.

[0419] Step 12:

[0420] The user (parent) approves or rejects the correction of the displayed inappropriate scene.

[0421] Step 13:

[0422] The server, upon receiving parental approval, queues a job to convert the inappropriate scenes to appropriate content.

[0423] Step 14:

[0424] The server launches the editing module to process the scene to be modified.

[0425] For example, frames depicting large amounts of blood could be converted to harmless versions or the scene could be removed entirely.

[0426] Step 15:

[0427] The server also edits the audio data to the appropriate content, maintaining overall synchronization.

[0428] Step 16:

[0429] The server generates a new video file after the editing is complete.

[0430] Step 17:

[0431] The server uploads the edited video to the user's dashboard and notifies them.

[0432] Step 18:

[0433] The user (parent) can check the edited video from the dashboard and play it on their device.

[0434] Step 19:

[0435] The user (child) can safely watch the edited video.

[0436] Example 1

[0437] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0438] There are many videos on the Internet that contain inappropriate content for children, making it difficult for parents to manage the content their children view. Editing out inappropriate scenes individually takes time and effort. To solve this problem, a system is needed that can automatically detect inappropriate content and edit it safely.

[0439] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0440] In this invention, the server includes a means for users to select or upload videos they wish to watch, a means for saving the video files in temporary storage and creating analysis jobs, and a means for transferring the video files to an artificial intelligence engine and requesting analysis. This makes it possible to automatically analyze videos containing inappropriate content and edit them in a safe manner.

[0441] "User" means an individual or entity that uses the System to select or upload videos that they wish to view.

[0442] A "video file" is a digital data file that contains visual and audio content.

[0443] "Temporary storage" is a digital storage device for temporarily storing video files until an analysis job is started.

[0444] An "analysis job" is a processing task that transfers a video file to an artificial intelligence engine and has its contents analyzed.

[0445] An "artificial intelligence engine" is a software or hardware system that analyzes video frames and audio data and runs algorithms to identify inappropriate language.

[0446] A "feature vector" is data that expresses each feature extracted from video frames or audio as a numerical value.

[0447] "Inappropriate content" refers to video scenes that contain violence, blood, or obscene content that is undesirable for children or general audiences.

[0448] A "timestamp" is data indicating the occurrence time of a frame containing an inappropriate expression.

[0449] A "scene tag" is a label that is assigned to identify information related to a particular scene.

[0450] "Parents" are guardians who are responsible for managing the content their children view.

[0451] The "overview panel" is an interface provided by the system, which is a dashboard for users to check the video analysis results and inappropriate scenes.

[0452] An "editing module" is a software or hardware system for changing detected inappropriate scenes into appropriate content.

[0453] "User Terminal" means a device used by a parent or user to access the system and view or manage videos.

[0454] A "new file" is a digital file generated containing the video modified by the editing module.

[0455] The present invention provides a system for providing an environment in which users can safely and appropriately watch videos they wish to watch. This system is realized using the following specific hardware and software.

[0456] 1. User video selection and upload

[0457] Users log in to the system's web application and select or upload the video they want to watch. To upload, the device sends the video file to the server using an HTTP POST request.

[0458] 2. Video transmission and storage to the server

[0459] The video file sent by the device is received by the server and saved in temporary storage (e.g., AWS S3 bucket, Google Cloud Storage), ensuring safe storage until the analysis job begins.

[0460] 3. Start video analysis

[0461] The server obtains the path of the saved video file, creates an analysis job, and adds it to the queue. A message broker such as RabbitMQ or Apache Kafka is used for the analysis job queue. The server then transfers the analysis job to the AI ​​engine and requests it to be analyzed.

[0462] 4. Video analysis using an AI engine

[0463] The AI ​​engine splits the video file into frames and analyzes the frames and audio data in parallel using deep learning models (e.g., YOLO, ResNet, and WaveNet for audio analysis). It extracts feature vectors from each frame and audio and determines whether the features represent inappropriate content (e.g., violence, blood, obscenity).

[0464] 5. Dashboard display of inappropriate scene information

[0465] The server generates data to display the analysis results from the AI ​​engine on the parent's dashboard. When the parent logs in, this data is displayed and the parent can check the timestamps of inappropriate scenes and scene tags.

[0466] 6. Parental approval of inappropriate scenes and corrections

[0467] Parents can check inappropriate scenes via the dashboard and approve corrections by sending a correction request to the server.

[0468] 7. Edit video with the editing module

[0469] After receiving parental approval, the server launches the editing module, which applies filters and replacements to inappropriate scenes based on the AI ​​engine's judgment. These modifications include changing the color of blood and muting obscene content.

[0470] 8. Video generation and delivery to devices

[0471] The modified video file is created as a new file and saved in temporary storage. The server sends a link to this file to the parent device and notifies them that editing is complete.

[0472] Specific examples

[0473] When uploading a "children's cartoon," the system analyzes the video and identifies violent scenes. If the parent reviews the details of the violent scene displayed on the dashboard and approves the edits, a scene in which a character slashes an enemy with a sword is changed to a soft stick, and the attack is deleted. A new file is generated and provided to the parent's device. An example prompt is, "Please explain the system for editing and providing children's videos to make them safe."

[0474] This invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. In addition, the automatic analysis and editing functions using AI make it possible to provide an appropriate video environment without any hassle.

[0475] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0476] Step 1:

[0477] User selects or uploads a video

[0478] A user logs in to a web application and selects or uploads a video they want to watch. When the user clicks the upload button for a video file, the device sends the video file to the server using an HTTP POST request. The input is the video file selected by the user, and the output is the transmission of the video file to the server. Specifically, the device reads the metadata of the video file, generates an upload request, and sends it to the server.

[0479] Step 2:

[0480] Sending and saving video files to the server

[0481] The video file sent from the device is received by the server. The server saves the received video file in temporary storage (e.g., AWS S3 bucket, Google Cloud Storage). The input is the video file sent from the device, and the output is the video file saved in temporary storage. Specifically, the server generates a path to save the video file and uploads the file to storage.

[0482] Step 3:

[0483] Starting a video analysis job

[0484] The server obtains the path of the saved video file, creates an analysis job, and adds it to a message queue. The input is the path of the video file saved in temporary storage, and the output is the job added to the analysis job queue. Specifically, the server generates a job ID and adds the job details, along with the path of the video file, to a queue (such as RabbitMQ or Apache Kafka).

[0485] Step 4:

[0486] Video analysis by AI engine

[0487] The server transfers the analysis job to the AI ​​engine and requests analysis. The AI ​​engine divides the video file into frames and analyzes the frame and audio data. The input is the analysis job details (video file path, job ID), and the output is the analysis result data (feature vector, timestamp, scene tag). Specifically, the AI ​​engine analyzes the images and audio using YOLO, ResNet, or WaveNet models, and extracts feature vectors from each frame and audio.

[0488] Step 5:

[0489] Dashboard display of inappropriate scene information

[0490] The server converts the analysis results received from the AI ​​engine into dashboard data and displays it on the parent dashboard. The input is the analysis result data of inappropriate scenes, and the output is the timestamp and scene tag displayed on the parent dashboard. Specifically, the server formats the data and sends it to the dashboard frontend using WebSocket or HTTP requests.

[0491] Step 6:

[0492] Parental review and approval of inappropriate scenes

[0493] Parents log in to the dashboard, check the inappropriate scenes, and approve their corrections. The input is a list of inappropriate scenes displayed on the dashboard, and the output is a request for approval of the corrections. Specifically, when a parent checks the scene details and thumbnails and clicks the "Approve Corrections" button, a request is sent to the server.

[0494] Step 7:

[0495] Editing module for video editing

[0496] The server launches the editing module after receiving approval for the correction from the parent. The editing module corrects the inappropriate scenes based on the AI ​​engine's judgment. The input is the correction approval request and details of the inappropriate scenes, and the output is the corrected video file. Specifically, the editing module detects inappropriate frames and applies filters and replacement processes to correct the content of the scenes.

[0497] Step 8:

[0498] Generate edited video and deliver it to the device

[0499] The server generates a new video file modified by the editing module and stores it in temporary storage. The input is the modified video data, and the output is a link to the new video file. Specifically, the server generates the video file and sends its path to the parent device.

[0500] Step 9:

[0501] Play and check the edited video

[0502] The parent uses the provided link to play the edited video on their device and confirm the content. The input is a link to the edited video file, and the output is the confirmed video playback. Specifically, the parent clicks the link to launch the video player and watches the edited video to confirm that it is safe for their child to watch.

[0503] (Application example 1)

[0504] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0505] With the recent spread of video streaming services, the safety of children's content has become a major issue. In particular, it is important to prevent videos containing inappropriate content from falling into the hands of children. However, it takes a great deal of time and effort for parents to manually review and edit the content of all videos. To solve this problem, an efficient and highly accurate automated system is needed.

[0506] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0507] In this invention, the server includes: means for a user to select or upload a video they wish to watch; a server for storing and analyzing video files; an artificial intelligence engine for extracting the analyzed video content as a feature vector and identifying scenes containing inappropriate content; means for tagging the inappropriate scenes identified by the artificial intelligence engine and notifying the user on their dashboard; means for the user to review the inappropriate scenes and approve their correction; an editing module for converting the inappropriate scenes into appropriate content; and means for generating edited videos and providing them to a content distribution service. This allows users to quickly and easily manage, edit, and provide safe videos for children without any hassle.

[0508] "Means for users to select or upload videos they wish to view" refers to a function that allows users to use their own devices to specify or upload videos they wish to view to the system.

[0509] A "server for storing and analyzing video files" is a computer device that temporarily stores uploaded video files and analyzes their contents using an artificial intelligence engine.

[0510] The "artificial intelligence engine for extracting the analyzed video content as feature vectors and identifying scenes containing inappropriate expressions" is a processing device equipped with an artificial intelligence algorithm for analyzing video frames and audio data and automatically detecting inappropriate scenes based on the resulting feature vectors.

[0511] "Means for tagging inappropriate scenes identified by the AI ​​engine and notifying the user on their dashboard" refers to a mechanism for tagging inappropriate scenes detected by the AI ​​engine and displaying that information on a dashboard where users can view it.

[0512] "Means for users to confirm inappropriate scenes and approve corrections" refers to a function that allows users to use the dashboard to confirm the content of inappropriate scenes and select and approve whether or not to correct the scenes.

[0513] The "editing module for converting inappropriate scenes into appropriate content" is a software module that has an editing function for converting scenes that are determined to be inappropriate into more appropriate content based on the user's approval.

[0514] "Means for generating edited video and providing it to a content distribution service" refers to a function for generating a new video file containing the modified scene and providing that video to users via a content distribution service.

[0515] The present invention relates to a system that provides a safe viewing environment for video content for children. In this system, the server, the terminal, and the user each play an important role.

[0516] First, a user uploads or selects a video they wish to watch using their device. The video file is then sent to the server and temporarily stored in storage. The server then transfers the received video file to an AI engine for analysis. The AI ​​engine analyzes the video frames and audio data to extract feature vectors. This allows it to identify scenes that contain inappropriate content, such as violent scenes or scenes containing inappropriate language.

[0517] The analysis results are extracted as timestamps and scene tags for inappropriate scenes identified by the AI ​​engine. This information is then posted to the user's dashboard. Parents or appropriate administrators can log in to the dashboard to review the inappropriate scenes and approve their corrections. For example, if a character uses a sword in a violent scene, the scene can be changed to a soft stick, or the scene itself can be replaced with different content.

[0518] Once approval is given, an editing module is launched on the server side and the inappropriate scenes are converted into appropriate content. This editing process is carried out automatically using artificial intelligence. The edited video is generated as a new file and provided to the user's device via the content distribution service. This ensures that children can watch videos safely.

[0519] The following hardware and software can be used to build the system. It is recommended that the server be equipped with storage, a CPU, and a GPU for storing and analyzing video files. Deep learning frameworks such as TensorFlow and PyTorch can be used as the artificial intelligence engine. Terminals include devices such as smartphones, tablets, and PCs.

[0520] As a concrete example, consider the case where a user uploads a children's cartoon. The system receives the video and identifies violent scenes as a result of analysis by an artificial intelligence engine. The dashboard displays timestamps and scene tags for violent scenes, and parents can click a correction button to edit the scene. For example, in a scene where a character slashes an enemy with a sword, the sword can be changed to a soft stick. In this way, an edited video is generated and provided to the user's device via a distribution service.

[0521] Examples of prompts include:

[0522] In order to analyze videos for inappropriate content and provide a safe viewing environment, please analyze and edit your videos according to the following requirements.

[0523] Identify any inappropriate scenes and provide their start and end times

[0524] If a scene needs editing, replace it with appropriate content.

[0525] In this way, users can quickly and easily manage, edit and provide safe videos for children without any hassle.

[0526] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0527] Step 1:

[0528] A user uses a terminal to upload the video they want to watch to the system. The input is the video file selected or uploaded by the user, and the output is the video file sent to the server, at which point the file is saved in temporary storage.

[0529] Step 2:

[0530] The server transfers the received video file to the AI ​​engine and begins analysis. The input is the video file stored on the server, and the output is the video data sent to the AI ​​engine. The server completes the video file transfer procedure and prepares for analysis.

[0531] Step 3:

[0532] The AI ​​engine analyzes video frames and audio data to extract feature vectors. The input is the video data transferred from the server, and the output is the analysis result: feature vectors, timestamps of inappropriate scenes, and scene tags. An AI model (e.g., TensorFlow or PyTorch) is used to analyze each frame in the video and perform data calculations to detect inappropriate content.

[0533] Step 4:

[0534] The server tags inappropriate scenes identified by the AI ​​engine and notifies the user's dashboard. The input is the analysis result from the AI ​​engine, and the output is the inappropriate scene information notified on the dashboard. The server reflects the tag information and timestamp on the dashboard.

[0535] Step 5:

[0536] The user logs in to the dashboard, checks the inappropriate scenes, and approves the corrections. The input is the inappropriate scene information displayed on the dashboard, and the output is a command to approve the corrections. The user's operation determines whether the inappropriate scenes need to be corrected.

[0537] Step 6:

[0538] The server starts the editing module based on the approved correction request. The input is a command from the user to approve the correction, and the output is the start of the editing module. The editing module automatically converts inappropriate scenes into appropriate ones.

[0539] Step 7:

[0540] The editing module converts inappropriate scenes into appropriate content and generates a new edited video file. The input is the inappropriate scene data handled by the editing module, and the output is the edited video file. The AI ​​model is used to correct the content of the scene and generate a new edited video.

[0541] Step 8:

[0542] The server saves the edited video file as a final file and provides it to the user's device through the content distribution service. The input is the edited video file, and the output is the video file provided to the user's device through the distribution service. Users can receive safe video content for children to watch.

[0543] This allows users to easily create a safe video viewing environment for children.

[0544] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0545] The present invention relates to a system for providing an environment in which users can watch videos they want to watch in a safe and appropriate manner. In particular, the present invention describes a system that combines an emotion engine that recognizes the user's emotions and efficiently identifies and corrects inappropriate scenes.

[0546] The system begins with the user selecting or uploading a video they want to watch. Once the user uploads the video, the device sends the video file to the server. The server then stores the received video file in temporary storage and starts the analysis job.

[0547] The server transfers the video file to an AI engine, which extracts feature vectors from the video frames and audio data. This allows a pre-trained model to determine whether the video contains inappropriate content. For example, scenes in which a character attacks another character with a weapon or contains a large amount of blood are identified as inappropriate content.

[0548] Furthermore, the present invention incorporates an emotion engine that collects emotional data from the user while watching. The emotion engine acquires emotional data such as the user's facial expressions and heart rate through cameras and sensors, and associates the data with inappropriate scenes. For example, if the user's heart rate temporarily rises or if a stressed reaction is observed in the user's facial expression, the emotion engine determines that the scene is likely to be inappropriate.

[0549] The server displays the timestamps and scene tags of inappropriate scenes identified by the AI ​​engine, as well as emotional data from the emotion engine, on the parent's dashboard. Parents log in to the dashboard and review the inappropriate scenes. If the parent approves the correction of a confirmed scene, the server launches an editing module to convert the inappropriate content into something appropriate. For example, in a scene depicting a large amount of blood, the color of the blood can be changed to a harmless liquid, or the scene itself can be replaced with softer content.

[0550] The edited video is then generated as a new file, which the server then provides to the parent's device. The parent then plays the edited video to ensure it is safe for their child to watch. The emotion engine also collects emotional data when the parent approves or rejects the edits, allowing the AI ​​engine to learn from this data to help identify inappropriate scenes in the future.

[0551] As a concrete example, consider the case where a parent uploads a "family movie" to the system. The system receives the video and analyzes it using AI and an emotion engine. The emotion engine collects user emotional data and determines violent or tragic scenes as inappropriate. Based on this information, the parent approves the scene modification on the dashboard, and the AI ​​engine converts the scene into appropriate content, resulting in a safe video.

[0552] In this way, this invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. Furthermore, by combining AI and an emotion engine, it is possible to identify inappropriate scenes with greater accuracy and make appropriate corrections.

[0553] The processing flow will be explained below.

[0554] Step 1:

[0555] Users log into the system's web application and select or upload the video they want to watch.

[0556] Step 2:

[0557] The terminal transmits the selected or uploaded video file to the server.

[0558] Step 3:

[0559] The server stores the received video file in temporary storage.

[0560] Step 4:

[0561] The server adds a job to the queue to start video analysis.

[0562] Step 5:

[0563] The server runs the analysis job and transfers the video file to the artificial intelligence engine.

[0564] Step 6:

[0565] The artificial intelligence engine analyzes video frames and extracts feature vectors for each frame.

[0566] Step 7:

[0567] The artificial intelligence engine analyzes the voice data and converts it into text data.

[0568] Step 8:

[0569] The artificial intelligence engine uses pre-trained models to determine whether a sentence contains inappropriate language.

[0570] Step 9:

[0571] The artificial intelligence engine will timestamp and tag inappropriate scenes.

[0572] Step 10:

[0573] The emotion engine collects emotional data while users watch videos, for example by analyzing their facial expressions using a camera and monitoring their heart rate with a heart rate sensor.

[0574] Step 11:

[0575] The emotion engine analyzes the collected emotion data and associates it with inappropriate scenes.

[0576] Step 12:

[0577] The server notifies the parent's dashboard of tagged inappropriate scenes and associated emotional data.

[0578] Step 13:

[0579] The user (parent) logs in to the dashboard and checks the inappropriate scenes and their emotion data.

[0580] Step 14:

[0581] The user (parent) approves or rejects the correction of the inappropriate scene.

[0582] Step 15:

[0583] The server, upon receiving parental approval, queues a job to convert the inappropriate scenes to appropriate content.

[0584] Step 16:

[0585] The server launches an editing module to process the scene to be altered, for example, converting a frame depicting large amounts of blood into a harmless liquid and replacing the scene with a different musical melody.

[0586] Step 17:

[0587] The server also edits the audio data to the appropriate content, maintaining overall synchronization.

[0588] Step 18:

[0589] The server generates a new video file after the editing is complete.

[0590] Step 19:

[0591] The server uploads the edited video to the user's dashboard and notifies them.

[0592] Step 20:

[0593] The user (parent) can check the edited video from the dashboard and play it on their device.

[0594] Step 21:

[0595] The user (child) can safely watch the edited video.

[0596] Step 22:

[0597] The emotion engine also collects emotional data when parents approve or reject modifications and feeds this back to the AI ​​engine.

[0598] Step 23:

[0599] The artificial intelligence engine uses the collected emotional data to optimize algorithms for identifying inappropriate scenes.

[0600] Example 2

[0601] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0602] Currently, when a scene containing inappropriate content is found in video content viewed by children, parents must manually detect and edit that scene. This places a significant burden on parents, and there are limitations to the accuracy of identifying and editing inappropriate scenes. Furthermore, there is no established method for effectively utilizing user emotion data to improve the accuracy of identifying inappropriate scenes. Therefore, there is a need for a system that provides a safe and appropriate video viewing environment and reduces the burden on parents.

[0603] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for selecting or uploading a video that a user wants to watch; means for saving and analyzing video files; means for extracting the content of the analyzed video as a feature vector and identifying scenes containing inappropriate expressions; means for collecting user emotion data and associating it with inappropriate scenes; means for tagging inappropriate scenes identified by the artificial intelligence engine and the emotion recognition engine and notifying them to a parent's dashboard; means for a parent to confirm and approve corrections of inappropriate scenes; means for converting inappropriate scenes into appropriate content; and means for generating edited videos and providing them to a user terminal. This enables automatic detection and editing of scenes containing inappropriate expressions, significantly reducing the burden on parents and providing an environment in which children can watch safe and appropriate videos.

[0604] "User" refers to the person who selects or uploads videos to the system that they wish to view.

[0605] "Video File" means a digital file containing video and audio data selected or uploaded by a User for viewing purposes.

[0606] A "server" refers to a computer system that stores video files and handles a series of processes such as analyzing and editing them.

[0607] "Feature vector" refers to vector data that quantifies specific attributes and features extracted from video frames and audio.

[0608] "Artificial intelligence engine" refers to a machine learning model that uses feature vectors to analyze videos and identify scenes containing inappropriate content.

[0609] An "emotion recognition engine" refers to a system that uses cameras and sensors to collect user emotional data and associates it with inappropriate scenes.

[0610] "Dashboard" refers to the user interface that allows parents to access the system to review information about inappropriate scenes and approve corrections.

[0611] "Editing Module" refers to software functionality for converting identified inappropriate scenes into appropriate content.

[0612] A "timestamp" refers to data that indicates the specific time at which a particular frame or scene appears in a video.

[0613] "Tags" refer to metadata used to classify and identify inappropriate scenes.

[0614] The present invention relates to a system for providing an environment in which a user can safely and appropriately view videos that the user desires to view. Detailed embodiments of the present invention will be described below.

[0615] First, the user selects or uploads the video they want to watch. The device sends the video file uploaded by the user to the server. The device can be a general personal computer or smartphone, and uses a file selection dialog that runs on a web browser.

[0616] Next, the server temporarily stores the received video file in a cloud storage service such as Amazon S3.

[0617] The server queues an analysis job for the stored video file and starts executing the job. The server sends the video file to an artificial intelligence engine (e.g., a TensorFlow model) that extracts feature vectors from the video frames and audio data.

[0618] The AI ​​engine uses feature vectors, specifically pre-trained machine learning models, to identify inappropriate scenes, such as scenes of violence or excessive gore.

[0619] Furthermore, the system collects user emotional data using cameras and sensors connected to the device. A standard webcam (e.g., Logitech C920) is used as the camera, and a heart rate monitor (e.g., Fitbit Charge 4) is used as the sensor. The Affectiva SDK is used as the emotion recognition engine. This allows the system to capture the user's facial expressions and heart rate in real time and associate them with inappropriate scenes.

[0620] The server displays the collected emotion data and timestamps and tags of inappropriate scenes on a dashboard. The dashboard is implemented using a web framework such as React.js. Parents can log in to the dashboard and check the inappropriate scenes.

[0621] If a parent approves the correction of an inappropriate scene, the server launches an editing module (e.g., FFmpeg command) that converts the inappropriate content into something appropriate, such as changing the color of blood or cutting the scene altogether.

[0622] Once the editing is complete, the video is generated as a new file, which is then provided to the device by the server. A link is displayed on the dashboard for parents to download the new video file. When parents click the link, the file is downloaded to the device.

[0623] Finally, parents can play the edited video on their device to confirm that it is safe for their child to watch. The emotion recognition engine also collects emotional data when parents approve or reject the edits, and uses this data for future learning on the server. This improves the accuracy of identifying inappropriate scenes in future videos.

[0624] As a concrete example, consider the case where a parent uploads a "family movie" to the system. The system receives the video and analyzes it using AI and an emotion recognition engine. For example, it uses a prompt such as, "Please detect whether this movie contains violent or tragic scenes." The emotion recognition engine collects the user's emotional data and determines that violent or tragic scenes are inappropriate. Based on this information, if the parent approves the scene modification on the dashboard, the AI ​​engine converts the scene into appropriate content, and a safe video is generated.

[0625] In this way, this invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. Furthermore, by combining AI and an emotion recognition engine, it is possible to identify inappropriate scenes with greater accuracy and make appropriate corrections.

[0626] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0627] Step 1:

[0628] The user selects or uploads the video they want to watch. The device retrieves the video file selected or uploaded by the user and sends it to the server using an HTTP POST request. The input is the video file selected by the user, and the output is the video file sent to the server.

[0629] Step 2:

[0630] The server saves the received video file in temporary storage. Specifically, it uploads the video file to cloud storage such as Amazon S3 and records the status of the saving completion in a log. The input is the video file sent to the server, and the output is the video file saved in the storage.

[0631] Step 3:

[0632] The server queues an analysis job for the stored video file and starts the job execution. The video file is sent to the artificial intelligence engine for analysis. The input is the video file stored in storage, and the output is the start status of the analysis job.

[0633] Step 4:

[0634] The server transfers the video file to an AI engine, which extracts feature vectors from the video frames and audio data. The AI ​​engine (e.g., a TensorFlow model) converts each video frame and audio data into a numerical vector and inputs it into a generative AI model. The input is the video file and its frames and audio data, and the output is the feature vector.

[0635] Step 5:

[0636] The AI ​​engine uses feature vectors to identify inappropriate scenes. Specifically, it applies a pre-trained model and uses a specific algorithm to detect frames containing inappropriate content. The input is the feature vectors, and the output is a list of inappropriate scenes with timestamps and tags.

[0637] Step 6:

[0638] The device uses an emotion recognition engine to collect user emotional data. It uses cameras and sensors connected to the device to capture the user's facial expressions and heart rate in real time. The collected data is sent to a server. The input is raw data from the video and heart rate sensors, and the output is emotional data collected in real time.

[0639] Step 7:

[0640] The server analyzes the collected emotion data and associates it with inappropriate scenes. The emotion recognition engine analyzes the collected data and identifies emotional changes in specific scenes. For example, it identifies when a person's heart rate spikes and associates this with inappropriate scenes. The input is emotion data, and the output is emotion data associated with inappropriate scenes.

[0641] Step 8:

[0642] The server displays the timestamps and tags of inappropriate scenes, as well as emotion data, on a dashboard. It uses React.js to display the information in a web interface that parents can access. The input is inappropriate scenes and emotion data, and the output is the data displayed on the dashboard.

[0643] Step 9:

[0644] The parent logs in to the dashboard and checks the inappropriate scenes. After checking, the parent "approves" or "rejects" the scene correction. The input is the data displayed on the dashboard, and the output is the parent's operation data.

[0645] Step 10:

[0646] The server starts the editing module and converts the content of the inappropriate scene into an appropriate one. Specifically, it uses FFmpeg commands to process the scene that needs to be edited. The input is the parent operation data and the inappropriate scene information, and the output is the edited scene.

[0647] Step 11:

[0648] The server generates a new video file containing the modified content. A new edited video file is created and saved to storage. The input is the edited scene, and the output is the new video file.

[0649] Step 12:

[0650] The server provides the generated new video file to the parent's device. A download link is generated on the dashboard, which the parent can access to download the new video. The input is the new video file, and the output is the download link.

[0651] Step 13:

[0652] The parent plays the edited video on the device and confirms that it is safe for the child to watch. The input is the new video file, and the output is the parent's confirmation result.

[0653] Step 14:

[0654] The emotion recognition engine also collects emotional data when parents approve or reject corrections, and the server uses this data for future learning. This improves the accuracy of identifying inappropriate scenes from the next time onwards. The input is the parent's emotional data, and the output is learning data.

[0655] (Application example 2)

[0656] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0657] In conventional digital content viewing systems, it is difficult for users to manually detect and correct inappropriate scenes in the content they are viewing. Simply deleting inappropriate scenes often detracts from the user's viewing experience. Furthermore, there is a lack of a mechanism for more precisely identifying inappropriate scenes using user emotional data. This makes it difficult for users to enjoy content safely and comfortably.

[0658] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0659] In this invention, the server includes a means for a user to select or upload digital content they wish to view, a means for saving and analyzing digital content files, and a machine learning engine that extracts the analyzed digital content content as a feature vector and identifies scenes containing inappropriate content. This makes it possible to identify inappropriate scenes using emotion data. The server also includes a means for tagging inappropriate scenes identified by the machine learning engine and notifying a parental interface, a means for a parent to confirm the inappropriate scenes and approve corrections, and an editing module for converting the inappropriate scenes into appropriate content, thereby providing a safe and comfortable viewing experience for users. Furthermore, the server includes a means for collecting user emotion data and associating it with inappropriate scenes using the emotion engine, and a means for collecting emotion data when a parent approves or rejects corrections and performing learning based on this data, making it possible to more precisely identify inappropriate scenes and make appropriate corrections.

[0660] "Digital content" means media, including video, audio, images, and text, that is stored and distributed in electronic form.

[0661] A "machine learning engine" is a program that contains algorithms that analyze data, automatically learn patterns and features without relying on human-defined rules, and make predictions and classifications.

[0662] "Parental interface" refers to a user interface used by a parent to notify the parent of an inappropriate scene, and to review and approve the correction.

[0663] "Inappropriate content" refers to scenes that contain content that may have an upsetting or harmful effect on viewers, such as violence, discrimination, or extreme tragedy.

[0664] The "emotion engine" is a system that uses cameras and sensors to acquire biometric information such as a user's facial expressions and heart rate, and analyzes their emotional state.

[0665] An "editing module" is a software component for converting the content of identified inappropriate scenes into appropriate content.

[0666] A "feature vector" is data that quantifies the content extracted from the frames and audio data of digital content.

[0667] An "algorithm" is a set of steps or a computational method for solving a particular problem.

[0668] A "timestamp" is data that indicates the time at which a particular frame or scene occurred.

[0669] "Learning means" is the process by which a machine learning engine improves its model based on collected data.

[0670] MODE FOR CARRYING OUT THE INVENTION

[0671] To implement this invention, the following configuration and operation procedures are required.

[0672] The system of the present invention includes a terminal for users to select or upload digital content they wish to view, and a server for storing and analyzing digital content files. The server is equipped with a machine learning engine and an emotion engine, and collects and analyzes user emotion data in real time.

[0673] First, a user uploads digital content to a device. The device then sends the digital content file to a server. The server stores the received digital content file in temporary storage and initiates an analysis job. The server then transfers the digital content file to a machine learning engine, which extracts feature vectors from video frames and audio data. This allows a pre-trained algorithm to determine whether the content contains inappropriate language.

[0674] At the same time, the user's device is equipped with a camera and sensors, which allow the emotion engine to capture emotional data such as the user's facial expressions and heart rate. When an inappropriate scene is identified, the scene's timestamp and scene tag are recorded. When an inappropriate scene is identified, this information is notified to the parental interface.

[0675] Parents can review inappropriate scenes and approve their edits via the interface. Once the edits are approved, the server launches an editing module to convert the inappropriate content into appropriate content. For example, in a scene that contains a large amount of blood, the color of the blood can be changed to a harmless liquid, or the scene itself can be replaced with softer content.

[0676] Once edits are complete, the digital content is generated as a new file, which the server then provides to the user's device. The user can then play the edited digital content for safe and comfortable viewing. The emotion engine also collects emotional data from parents when they approve or reject the edits, which the machine learning engine can use to learn from to help identify inappropriate scenes in the future.

[0677] As a concrete example, consider the case where a children's animated content is uploaded. The system analyzes every frame and detects violent or tragic scenes. Parents can then review the content of the scenes on the dashboard and approve the edits. The edited video can then be safely shown to children.

[0678] An example of a prompt statement can be specified as follows:

[0679] Given the following text, generate a step-by-step guide to detecting and correcting inappropriate scenes in a children's cartoon, including collecting emotion data and analyzing the scene using AI.

[0680] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0681] Step 1:

[0682] The user selects or uploads digital content on the device.

[0683] Input: Digital content files (video files, audio files, etc.)

[0684] Output: The digital content file is saved to the server.

[0685] Specific operation: The user selects digital content through the terminal interface and clicks the upload button, after which the file is sent to the server.

[0686] Step 2:

[0687] The server stores the digital content file in temporary storage and starts the analysis job.

[0688] Input: Digital content file submitted by the user

[0689] Output: The digital content file is saved to temporary storage and analysis begins.

[0690] What happens: The server receives the digital content file and automatically stores it in temporary storage, then prepares it for analysis script execution.

[0691] Step 3:

[0692] The server transfers the digital content files to a machine learning engine, which extracts feature vectors for the video frames and audio data.

[0693] Input: Digital content file

[0694] Output: Feature vector (frame data, numerical representation of audio data)

[0695] Specific operation: The server transfers the digital content file to the machine learning engine, which processes the file to extract feature vectors from the frame data and audio data.

[0696] Step 4:

[0697] A machine learning engine uses feature vectors to detect scenes containing inappropriate language and record them with a timestamp.

[0698] Input: feature vector

[0699] Output: timestamp and tag data for the incorrect scene

[0700] How it works: The machine learning engine uses trained algorithms to analyze feature vectors and identify scenes that contain inappropriate content, such as violence or discrimination.

[0701] Step 5:

[0702] The device collects user emotional data in real time through cameras and sensors, and the emotion engine analyzes it.

[0703] Input: User's facial expression data, heart rate data

[0704] Output: Emotional data (stress response, heart rate variability, etc.)

[0705] Specific operation: The device's camera and sensors collect the user's facial expressions and biometric information, which are then sent to the emotion engine for analysis.

[0706] Step 6:

[0707] The emotion engine associates it with an inappropriate scene and sends it to the server.

[0708] Input: Emotion data, timestamps and tag data of inappropriate scenes

[0709] Output: Emotion data and association information for inappropriate scenes

[0710] Specific operation: The emotion engine matches the user's emotion data with the timestamp of the inappropriate scene based on the log.

[0711] Step 7:

[0712] The parental interface will be notified of inappropriate scenes and asked for approval to make corrections.

[0713] Input: timestamps and tag data of inappropriate scenes, emotion data

[0714] Output: Parental notification, request for approval of corrections

[0715] Specific behavior: The server displays detailed information about the inappropriate scene in the parental interface and asks the parent to approve or reject the correction.

[0716] Step 8:

[0717] After receiving parental approval, the server launches an editing module and converts the inappropriate scenes into appropriate content.

[0718] Input: Parental edit approval, timestamp and tag data for inappropriate scenes

[0719] Output: Edited scene data

[0720] Specific operation: The server's editing module converts the specified inappropriate scene into an appropriate expression and generates new scene data.

[0721] Step 9:

[0722] The edited digital content is generated as a new file and provided to the user's terminal.

[0723] Input: Edited scene data, original digital content excluding inappropriate scenes

[0724] Output: Edited digital content file

[0725] Specific operation: The server integrates the edited scene data into the entire original content, generates a new digital content file, and sends it to the user's device.

[0726] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0727] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0728] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0729] [Third embodiment]

[0730] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0731] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0732] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0733] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0734] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0735] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0736] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0737] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0738] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0739] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0740] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0741] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0742] The present invention relates to a system for providing an environment in which users can safely and appropriately watch videos they wish to watch. As shown below, the processing contents of a program in which the server, terminal, and user each play an important role are described.

[0743] First, a user logs in to the system's web application and selects or uploads the video they want to watch. At this time, the device sends the selected video file to the server. The server stores the received video file in temporary storage and starts the analysis job.

[0744] The server then transfers the video file to an AI engine for analysis. The AI ​​engine analyzes the video frames and audio data to extract feature vectors. This determines whether the video contains inappropriate content. For example, scenes in which a character attacks another character with a weapon or scenes containing large amounts of blood are identified as inappropriate content.

[0745] The server then displays the timestamps and scene tags of inappropriate scenes identified by the AI ​​engine on the parent's dashboard. Parents can log in to the dashboard and review the inappropriate scenes. If the parent approves the modification of the inappropriate scene, the server launches an editing module to convert the inappropriate content into an appropriate one. For example, in a scene depicting a large amount of blood, the color of the blood can be changed to a harmless liquid, or the scene itself can be replaced with softer content.

[0746] Once the editing is complete, the video is generated as a new file, which the server provides to the parent's device. The parent can then play the edited video on the device and confirm that the child is safe to watch.

[0747] As a concrete example, consider the case where a parent uploads a "children's cartoon" to the system. The system receives the video and identifies violent scenes as a result of analysis. The violent scenes are displayed on the dashboard, and when the parent clicks the approval button, the scenes are modified. For example, in a scene where a character slashes an enemy with a sword, the sword is changed to a soft stick and the depiction of the attack itself is deleted. This edited video is then provided to the parent's device, allowing the child to watch it safely.

[0748] This invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. In addition, the automatic analysis and editing functions using AI make it possible to provide an appropriate video environment without any hassle.

[0749] The processing flow will be explained below.

[0750] Step 1:

[0751] Users log into the system's web application and select or upload the video they want to watch.

[0752] Step 2:

[0753] The terminal transmits the selected or uploaded video file to the server.

[0754] Step 3:

[0755] The server stores the received video file in temporary storage.

[0756] Step 4:

[0757] The server adds a job to the queue to start video analysis.

[0758] Step 5:

[0759] The server runs the analysis job and transfers the video file to the artificial intelligence engine.

[0760] Step 6:

[0761] The artificial intelligence engine analyzes video frames and extracts feature vectors for each frame.

[0762] Step 7:

[0763] The artificial intelligence engine also analyzes voice data and converts it into text data.

[0764] Step 8:

[0765] The artificial intelligence engine uses pre-trained models to determine whether or not there is inappropriate language.

[0766] Step 9:

[0767] The artificial intelligence engine then timestamps and tags identified inappropriate scenes.

[0768] Step 10:

[0769] The server notifies the parent's dashboard of the tagged inappropriate scenes.

[0770] Step 11:

[0771] The user (parent) logs in to the dashboard and checks for inappropriate scenes.

[0772] Step 12:

[0773] The user (parent) approves or rejects the correction of the displayed inappropriate scene.

[0774] Step 13:

[0775] The server, upon receiving parental approval, queues a job to convert the inappropriate scenes to appropriate content.

[0776] Step 14:

[0777] The server launches the editing module to process the scene to be modified.

[0778] For example, frames depicting large amounts of blood could be converted to harmless versions or the scene could be removed entirely.

[0779] Step 15:

[0780] The server also edits the audio data to the appropriate content, maintaining overall synchronization.

[0781] Step 16:

[0782] The server generates a new video file after the editing is complete.

[0783] Step 17:

[0784] The server uploads the edited video to the user's dashboard and notifies them.

[0785] Step 18:

[0786] The user (parent) can check the edited video from the dashboard and play it on their device.

[0787] Step 19:

[0788] The user (child) can safely watch the edited video.

[0789] Example 1

[0790] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0791] There are many videos on the Internet that contain inappropriate content for children, making it difficult for parents to manage the content their children view. Editing out inappropriate scenes individually takes time and effort. To solve this problem, a system is needed that can automatically detect inappropriate content and edit it safely.

[0792] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0793] In this invention, the server includes a means for users to select or upload videos they wish to watch, a means for saving the video files in temporary storage and creating analysis jobs, and a means for transferring the video files to an artificial intelligence engine and requesting analysis. This makes it possible to automatically analyze videos containing inappropriate content and edit them in a safe manner.

[0794] "User" means an individual or entity that uses the System to select or upload videos that they wish to view.

[0795] A "video file" is a digital data file that contains visual and audio content.

[0796] "Temporary storage" is a digital storage device for temporarily storing video files until an analysis job is started.

[0797] An "analysis job" is a processing task that transfers a video file to an artificial intelligence engine and has its contents analyzed.

[0798] An "artificial intelligence engine" is a software or hardware system that analyzes video frames and audio data and runs algorithms to identify inappropriate language.

[0799] A "feature vector" is data that expresses each feature extracted from video frames or audio as a numerical value.

[0800] "Inappropriate content" refers to video scenes that contain violence, blood, or obscene content that is undesirable for children or general audiences.

[0801] A "timestamp" is data indicating the occurrence time of a frame containing an inappropriate expression.

[0802] A "scene tag" is a label that is assigned to identify information related to a particular scene.

[0803] "Parents" are guardians who are responsible for managing the content their children view.

[0804] The "overview panel" is an interface provided by the system, which is a dashboard for users to check the video analysis results and inappropriate scenes.

[0805] An "editing module" is a software or hardware system for changing detected inappropriate scenes into appropriate content.

[0806] "User Terminal" means a device used by a parent or user to access the system and view or manage videos.

[0807] A "new file" is a digital file generated containing the video modified by the editing module.

[0808] The present invention provides a system for providing an environment in which users can safely and appropriately watch videos they wish to watch. This system is realized using the following specific hardware and software.

[0809] 1. User video selection and upload

[0810] Users log in to the system's web application and select or upload the video they want to watch. To upload, the device sends the video file to the server using an HTTP POST request.

[0811] 2. Video transmission and storage to the server

[0812] The video file sent by the device is received by the server and saved in temporary storage (e.g., AWS S3 bucket, Google Cloud Storage), ensuring safe storage until the analysis job begins.

[0813] 3. Start video analysis

[0814] The server obtains the path of the saved video file, creates an analysis job, and adds it to the queue. A message broker such as RabbitMQ or Apache Kafka is used for the analysis job queue. The server then transfers the analysis job to the AI ​​engine and requests it to be analyzed.

[0815] 4. Video analysis using an AI engine

[0816] The AI ​​engine splits the video file into frames and analyzes the frames and audio data in parallel using deep learning models (e.g., YOLO, ResNet, and WaveNet for audio analysis). It extracts feature vectors from each frame and audio and determines whether the features represent inappropriate content (e.g., violence, blood, obscenity).

[0817] 5. Dashboard display of inappropriate scene information

[0818] The server generates data to display the analysis results from the AI ​​engine on the parent's dashboard. When the parent logs in, this data is displayed and the parent can check the timestamps of inappropriate scenes and scene tags.

[0819] 6. Parental approval of inappropriate scenes and corrections

[0820] Parents can check inappropriate scenes via the dashboard and approve corrections by sending a correction request to the server.

[0821] 7. Edit video with the editing module

[0822] After receiving parental approval, the server launches the editing module, which applies filters and replacements to inappropriate scenes based on the AI ​​engine's judgment. These modifications include changing the color of blood and muting obscene content.

[0823] 8. Video generation and delivery to devices

[0824] The modified video file is created as a new file and saved in temporary storage. The server sends a link to this file to the parent device and notifies them that editing is complete.

[0825] Specific examples

[0826] When uploading a "children's cartoon," the system analyzes the video and identifies violent scenes. If the parent reviews the details of the violent scene displayed on the dashboard and approves the edits, a scene in which a character slashes an enemy with a sword is changed to a soft stick, and the attack is deleted. A new file is generated and provided to the parent's device. An example prompt is, "Please explain the system for editing and providing children's videos to make them safe."

[0827] This invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. In addition, the automatic analysis and editing functions using AI make it possible to provide an appropriate video environment without any hassle.

[0828] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0829] Step 1:

[0830] User selects or uploads a video

[0831] A user logs in to a web application and selects or uploads a video they want to watch. When the user clicks the upload button for a video file, the device sends the video file to the server using an HTTP POST request. The input is the video file selected by the user, and the output is the transmission of the video file to the server. Specifically, the device reads the metadata of the video file, generates an upload request, and sends it to the server.

[0832] Step 2:

[0833] Sending and saving video files to the server

[0834] The video file sent from the device is received by the server. The server saves the received video file in temporary storage (e.g., AWS S3 bucket, Google Cloud Storage). The input is the video file sent from the device, and the output is the video file saved in temporary storage. Specifically, the server generates a path to save the video file and uploads the file to storage.

[0835] Step 3:

[0836] Starting a video analysis job

[0837] The server obtains the path of the saved video file, creates an analysis job, and adds it to a message queue. The input is the path of the video file saved in temporary storage, and the output is the job added to the analysis job queue. Specifically, the server generates a job ID and adds the job details, along with the path of the video file, to a queue (such as RabbitMQ or Apache Kafka).

[0838] Step 4:

[0839] Video analysis by AI engine

[0840] The server transfers the analysis job to the AI ​​engine and requests analysis. The AI ​​engine divides the video file into frames and analyzes the frame and audio data. The input is the analysis job details (video file path, job ID), and the output is the analysis result data (feature vector, timestamp, scene tag). Specifically, the AI ​​engine analyzes the images and audio using YOLO, ResNet, or WaveNet models, and extracts feature vectors from each frame and audio.

[0841] Step 5:

[0842] Dashboard display of inappropriate scene information

[0843] The server converts the analysis results received from the AI ​​engine into dashboard data and displays it on the parent dashboard. The input is the analysis result data of inappropriate scenes, and the output is the timestamp and scene tag displayed on the parent dashboard. Specifically, the server formats the data and sends it to the dashboard frontend using WebSocket or HTTP requests.

[0844] Step 6:

[0845] Parental review and approval of inappropriate scenes

[0846] Parents log in to the dashboard, check the inappropriate scenes, and approve their corrections. The input is a list of inappropriate scenes displayed on the dashboard, and the output is a request for approval of the corrections. Specifically, when a parent checks the scene details and thumbnails and clicks the "Approve Corrections" button, a request is sent to the server.

[0847] Step 7:

[0848] Editing module for video editing

[0849] The server launches the editing module after receiving approval for the correction from the parent. The editing module corrects the inappropriate scenes based on the AI ​​engine's judgment. The input is the correction approval request and details of the inappropriate scenes, and the output is the corrected video file. Specifically, the editing module detects inappropriate frames and applies filters and replacement processes to correct the content of the scenes.

[0850] Step 8:

[0851] Generate edited video and deliver it to the device

[0852] The server generates a new video file modified by the editing module and stores it in temporary storage. The input is the modified video data, and the output is a link to the new video file. Specifically, the server generates the video file and sends its path to the parent device.

[0853] Step 9:

[0854] Play and check the edited video

[0855] The parent uses the provided link to play the edited video on their device and confirm the content. The input is a link to the edited video file, and the output is the confirmed video playback. Specifically, the parent clicks the link to launch the video player and watches the edited video to confirm that it is safe for their child to watch.

[0856] (Application example 1)

[0857] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0858] With the recent spread of video streaming services, the safety of children's content has become a major issue. In particular, it is important to prevent videos containing inappropriate content from falling into the hands of children. However, it takes a great deal of time and effort for parents to manually review and edit the content of all videos. To solve this problem, an efficient and highly accurate automated system is needed.

[0859] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0860] In this invention, the server includes: means for a user to select or upload a video they wish to watch; a server for storing and analyzing video files; an artificial intelligence engine for extracting the analyzed video content as a feature vector and identifying scenes containing inappropriate content; means for tagging the inappropriate scenes identified by the artificial intelligence engine and notifying the user on their dashboard; means for the user to review the inappropriate scenes and approve their correction; an editing module for converting the inappropriate scenes into appropriate content; and means for generating edited videos and providing them to a content distribution service. This allows users to quickly and easily manage, edit, and provide safe videos for children without any hassle.

[0861] "Means for users to select or upload videos they wish to view" refers to a function that allows users to use their own devices to specify or upload videos they wish to view to the system.

[0862] A "server for storing and analyzing video files" is a computer device that temporarily stores uploaded video files and analyzes their contents using an artificial intelligence engine.

[0863] The "artificial intelligence engine for extracting the analyzed video content as feature vectors and identifying scenes containing inappropriate expressions" is a processing device equipped with an artificial intelligence algorithm for analyzing video frames and audio data and automatically detecting inappropriate scenes based on the resulting feature vectors.

[0864] "Means for tagging inappropriate scenes identified by the AI ​​engine and notifying the user on their dashboard" refers to a mechanism for tagging inappropriate scenes detected by the AI ​​engine and displaying that information on a dashboard where users can view it.

[0865] "Means for users to confirm inappropriate scenes and approve corrections" refers to a function that allows users to use the dashboard to confirm the content of inappropriate scenes and select and approve whether or not to correct the scenes.

[0866] The "editing module for converting inappropriate scenes into appropriate content" is a software module that has an editing function for converting scenes that are determined to be inappropriate into more appropriate content based on the user's approval.

[0867] "Means for generating edited video and providing it to a content distribution service" refers to a function for generating a new video file containing the modified scene and providing that video to users via a content distribution service.

[0868] The present invention relates to a system that provides a safe viewing environment for video content for children. In this system, the server, the terminal, and the user each play an important role.

[0869] First, a user uploads or selects a video they wish to watch using their device. The video file is then sent to the server and temporarily stored in storage. The server then transfers the received video file to an AI engine for analysis. The AI ​​engine analyzes the video frames and audio data to extract feature vectors. This allows it to identify scenes that contain inappropriate content, such as violent scenes or scenes containing inappropriate language.

[0870] The analysis results are extracted as timestamps and scene tags for inappropriate scenes identified by the AI ​​engine. This information is then posted to the user's dashboard. Parents or appropriate administrators can log in to the dashboard to review the inappropriate scenes and approve their corrections. For example, if a character uses a sword in a violent scene, the scene can be changed to a soft stick, or the scene itself can be replaced with different content.

[0871] Once approval is given, an editing module is launched on the server side and the inappropriate scenes are converted into appropriate content. This editing process is carried out automatically using artificial intelligence. The edited video is generated as a new file and provided to the user's device via the content distribution service. This ensures that children can watch videos safely.

[0872] The following hardware and software can be used to build the system. It is recommended that the server be equipped with storage, a CPU, and a GPU for storing and analyzing video files. Deep learning frameworks such as TensorFlow and PyTorch can be used as the artificial intelligence engine. Terminals include devices such as smartphones, tablets, and PCs.

[0873] As a concrete example, consider the case where a user uploads a children's cartoon. The system receives the video and identifies violent scenes as a result of analysis by an artificial intelligence engine. The dashboard displays timestamps and scene tags for violent scenes, and parents can click a correction button to edit the scene. For example, in a scene where a character slashes an enemy with a sword, the sword can be changed to a soft stick. In this way, an edited video is generated and provided to the user's device via a distribution service.

[0874] Examples of prompts include:

[0875] In order to analyze videos for inappropriate content and provide a safe viewing environment, please analyze and edit your videos according to the following requirements.

[0876] Identify any inappropriate scenes and provide their start and end times

[0877] If a scene needs editing, replace it with appropriate content.

[0878] In this way, users can quickly and easily manage, edit and provide safe videos for children without any hassle.

[0879] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0880] Step 1:

[0881] A user uses a terminal to upload the video they want to watch to the system. The input is the video file selected or uploaded by the user, and the output is the video file sent to the server, at which point the file is saved in temporary storage.

[0882] Step 2:

[0883] The server transfers the received video file to the AI ​​engine and begins analysis. The input is the video file stored on the server, and the output is the video data sent to the AI ​​engine. The server completes the video file transfer procedure and prepares for analysis.

[0884] Step 3:

[0885] The AI ​​engine analyzes video frames and audio data to extract feature vectors. The input is the video data transferred from the server, and the output is the analysis result: feature vectors, timestamps of inappropriate scenes, and scene tags. An AI model (e.g., TensorFlow or PyTorch) is used to analyze each frame in the video and perform data calculations to detect inappropriate content.

[0886] Step 4:

[0887] The server tags inappropriate scenes identified by the AI ​​engine and notifies the user's dashboard. The input is the analysis result from the AI ​​engine, and the output is the inappropriate scene information notified on the dashboard. The server reflects the tag information and timestamp on the dashboard.

[0888] Step 5:

[0889] The user logs in to the dashboard, checks the inappropriate scenes, and approves the corrections. The input is the inappropriate scene information displayed on the dashboard, and the output is a command to approve the corrections. The user's operation determines whether the inappropriate scenes need to be corrected.

[0890] Step 6:

[0891] The server starts the editing module based on the approved correction request. The input is a command from the user to approve the correction, and the output is the start of the editing module. The editing module automatically converts inappropriate scenes into appropriate ones.

[0892] Step 7:

[0893] The editing module converts inappropriate scenes into appropriate content and generates a new edited video file. The input is the inappropriate scene data handled by the editing module, and the output is the edited video file. The AI ​​model is used to correct the content of the scene and generate a new edited video.

[0894] Step 8:

[0895] The server saves the edited video file as a final file and provides it to the user's device through the content distribution service. The input is the edited video file, and the output is the video file provided to the user's device through the distribution service. Users can receive safe video content for children to watch.

[0896] This allows users to easily create a safe video viewing environment for children.

[0897] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0898] The present invention relates to a system for providing an environment in which users can watch videos they want to watch in a safe and appropriate manner. In particular, the present invention describes a system that combines an emotion engine that recognizes the user's emotions and efficiently identifies and corrects inappropriate scenes.

[0899] The system begins with the user selecting or uploading a video they want to watch. Once the user uploads the video, the device sends the video file to the server. The server then stores the received video file in temporary storage and starts the analysis job.

[0900] The server transfers the video file to an AI engine, which extracts feature vectors from the video frames and audio data. This allows a pre-trained model to determine whether the video contains inappropriate content. For example, scenes in which a character attacks another character with a weapon or contains a large amount of blood are identified as inappropriate content.

[0901] Furthermore, the present invention incorporates an emotion engine that collects emotional data from the user while watching. The emotion engine acquires emotional data such as the user's facial expressions and heart rate through cameras and sensors, and associates the data with inappropriate scenes. For example, if the user's heart rate temporarily rises or if a stressed reaction is observed in the user's facial expression, the emotion engine determines that the scene is likely to be inappropriate.

[0902] The server displays the timestamps and scene tags of inappropriate scenes identified by the AI ​​engine, as well as emotional data from the emotion engine, on the parent's dashboard. Parents log in to the dashboard and review the inappropriate scenes. If the parent approves the correction of a confirmed scene, the server launches an editing module to convert the inappropriate content into something appropriate. For example, in a scene depicting a large amount of blood, the color of the blood can be changed to a harmless liquid, or the scene itself can be replaced with softer content.

[0903] The edited video is then generated as a new file, which the server then provides to the parent's device. The parent then plays the edited video to ensure it is safe for their child to watch. The emotion engine also collects emotional data when the parent approves or rejects the edits, allowing the AI ​​engine to learn from this data to help identify inappropriate scenes in the future.

[0904] As a concrete example, consider the case where a parent uploads a "family movie" to the system. The system receives the video and analyzes it using AI and an emotion engine. The emotion engine collects user emotional data and determines violent or tragic scenes as inappropriate. Based on this information, the parent approves the scene modification on the dashboard, and the AI ​​engine converts the scene into appropriate content, resulting in a safe video.

[0905] In this way, this invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. Furthermore, by combining AI and an emotion engine, it is possible to identify inappropriate scenes with greater accuracy and make appropriate corrections.

[0906] The processing flow will be explained below.

[0907] Step 1:

[0908] Users log into the system's web application and select or upload the video they want to watch.

[0909] Step 2:

[0910] The terminal transmits the selected or uploaded video file to the server.

[0911] Step 3:

[0912] The server stores the received video file in temporary storage.

[0913] Step 4:

[0914] The server adds a job to the queue to start video analysis.

[0915] Step 5:

[0916] The server runs the analysis job and transfers the video file to the artificial intelligence engine.

[0917] Step 6:

[0918] The artificial intelligence engine analyzes video frames and extracts feature vectors for each frame.

[0919] Step 7:

[0920] The artificial intelligence engine analyzes the voice data and converts it into text data.

[0921] Step 8:

[0922] The artificial intelligence engine uses pre-trained models to determine whether a sentence contains inappropriate language.

[0923] Step 9:

[0924] The artificial intelligence engine will timestamp and tag inappropriate scenes.

[0925] Step 10:

[0926] The emotion engine collects emotional data while users watch videos, for example by analyzing their facial expressions using a camera and monitoring their heart rate with a heart rate sensor.

[0927] Step 11:

[0928] The emotion engine analyzes the collected emotion data and associates it with inappropriate scenes.

[0929] Step 12:

[0930] The server notifies the parent's dashboard of tagged inappropriate scenes and associated emotional data.

[0931] Step 13:

[0932] The user (parent) logs in to the dashboard and checks the inappropriate scenes and their emotion data.

[0933] Step 14:

[0934] The user (parent) approves or rejects the correction of the inappropriate scene.

[0935] Step 15:

[0936] The server, upon receiving parental approval, queues a job to convert the inappropriate scenes to appropriate content.

[0937] Step 16:

[0938] The server launches an editing module to process the scene to be altered, for example, converting a frame depicting large amounts of blood into a harmless liquid and replacing the scene with a different musical melody.

[0939] Step 17:

[0940] The server also edits the audio data to the appropriate content, maintaining overall synchronization.

[0941] Step 18:

[0942] The server generates a new video file after the editing is complete.

[0943] Step 19:

[0944] The server uploads the edited video to the user's dashboard and notifies them.

[0945] Step 20:

[0946] The user (parent) can check the edited video from the dashboard and play it on their device.

[0947] Step 21:

[0948] The user (child) can safely watch the edited video.

[0949] Step 22:

[0950] The emotion engine also collects emotional data when parents approve or reject modifications and feeds this back to the AI ​​engine.

[0951] Step 23:

[0952] The artificial intelligence engine uses the collected emotional data to optimize algorithms for identifying inappropriate scenes.

[0953] Example 2

[0954] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0955] Currently, when a scene containing inappropriate content is found in video content viewed by children, parents must manually detect and edit that scene. This places a significant burden on parents, and there are limitations to the accuracy of identifying and editing inappropriate scenes. Furthermore, there is no established method for effectively utilizing user emotion data to improve the accuracy of identifying inappropriate scenes. Therefore, there is a need for a system that provides a safe and appropriate video viewing environment and reduces the burden on parents.

[0956] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for selecting or uploading a video that a user wants to watch; means for saving and analyzing video files; means for extracting the content of the analyzed video as a feature vector and identifying scenes containing inappropriate expressions; means for collecting user emotion data and associating it with inappropriate scenes; means for tagging inappropriate scenes identified by the artificial intelligence engine and the emotion recognition engine and notifying them to a parent's dashboard; means for a parent to confirm and approve corrections of inappropriate scenes; means for converting inappropriate scenes into appropriate content; and means for generating edited videos and providing them to a user terminal. This enables automatic detection and editing of scenes containing inappropriate expressions, significantly reducing the burden on parents and providing an environment in which children can watch safe and appropriate videos.

[0957] "User" refers to the person who selects or uploads videos to the system that they wish to view.

[0958] "Video File" means a digital file containing video and audio data selected or uploaded by a User for viewing purposes.

[0959] A "server" refers to a computer system that stores video files and handles a series of processes such as analyzing and editing them.

[0960] "Feature vector" refers to vector data that quantifies specific attributes and features extracted from video frames and audio.

[0961] "Artificial intelligence engine" refers to a machine learning model that uses feature vectors to analyze videos and identify scenes containing inappropriate content.

[0962] An "emotion recognition engine" refers to a system that uses cameras and sensors to collect user emotional data and associates it with inappropriate scenes.

[0963] "Dashboard" refers to the user interface that allows parents to access the system to review information about inappropriate scenes and approve corrections.

[0964] "Editing Module" refers to software functionality for converting identified inappropriate scenes into appropriate content.

[0965] A "timestamp" refers to data that indicates the specific time at which a particular frame or scene appears in a video.

[0966] "Tags" refer to metadata used to classify and identify inappropriate scenes.

[0967] The present invention relates to a system for providing an environment in which a user can safely and appropriately view videos that the user desires to view. Detailed embodiments of the present invention will be described below.

[0968] First, the user selects or uploads the video they want to watch. The device sends the video file uploaded by the user to the server. The device can be a general personal computer or smartphone, and uses a file selection dialog that runs on a web browser.

[0969] Next, the server temporarily stores the received video file in a cloud storage service such as Amazon S3.

[0970] The server queues an analysis job for the stored video file and starts executing the job. The server sends the video file to an artificial intelligence engine (e.g., a TensorFlow model) that extracts feature vectors from the video frames and audio data.

[0971] The AI ​​engine uses feature vectors, specifically pre-trained machine learning models, to identify inappropriate scenes, such as scenes of violence or excessive gore.

[0972] Furthermore, the system collects user emotional data using cameras and sensors connected to the device. A standard webcam (e.g., Logitech C920) is used as the camera, and a heart rate monitor (e.g., Fitbit Charge 4) is used as the sensor. The Affectiva SDK is used as the emotion recognition engine. This allows the system to capture the user's facial expressions and heart rate in real time and associate them with inappropriate scenes.

[0973] The server displays the collected emotion data and timestamps and tags of inappropriate scenes on a dashboard. The dashboard is implemented using a web framework such as React.js. Parents can log in to the dashboard and check the inappropriate scenes.

[0974] If a parent approves the correction of an inappropriate scene, the server launches an editing module (e.g., FFmpeg command) that converts the inappropriate content into something appropriate, such as changing the color of blood or cutting the scene altogether.

[0975] Once the editing is complete, the video is generated as a new file, which is then provided to the device by the server. A link is displayed on the dashboard for parents to download the new video file. When parents click the link, the file is downloaded to the device.

[0976] Finally, parents can play the edited video on their device to confirm that it is safe for their child to watch. The emotion recognition engine also collects emotional data when parents approve or reject the edits, and uses this data for future learning on the server. This improves the accuracy of identifying inappropriate scenes in future videos.

[0977] As a concrete example, consider the case where a parent uploads a "family movie" to the system. The system receives the video and analyzes it using AI and an emotion recognition engine. For example, it uses a prompt such as, "Please detect whether this movie contains violent or tragic scenes." The emotion recognition engine collects the user's emotional data and determines that violent or tragic scenes are inappropriate. Based on this information, if the parent approves the scene modification on the dashboard, the AI ​​engine converts the scene into appropriate content, and a safe video is generated.

[0978] In this way, this invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. Furthermore, by combining AI and an emotion recognition engine, it is possible to identify inappropriate scenes with greater accuracy and make appropriate corrections.

[0979] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0980] Step 1:

[0981] The user selects or uploads the video they want to watch. The device retrieves the video file selected or uploaded by the user and sends it to the server using an HTTP POST request. The input is the video file selected by the user, and the output is the video file sent to the server.

[0982] Step 2:

[0983] The server saves the received video file in temporary storage. Specifically, it uploads the video file to cloud storage such as Amazon S3 and records the status of the saving completion in a log. The input is the video file sent to the server, and the output is the video file saved in the storage.

[0984] Step 3:

[0985] The server queues an analysis job for the stored video file and starts the job execution. The video file is sent to the artificial intelligence engine for analysis. The input is the video file stored in storage, and the output is the start status of the analysis job.

[0986] Step 4:

[0987] The server transfers the video file to an AI engine, which extracts feature vectors from the video frames and audio data. The AI ​​engine (e.g., a TensorFlow model) converts each video frame and audio data into a numerical vector and inputs it into a generative AI model. The input is the video file and its frames and audio data, and the output is the feature vector.

[0988] Step 5:

[0989] The AI ​​engine uses feature vectors to identify inappropriate scenes. Specifically, it applies a pre-trained model and uses a specific algorithm to detect frames containing inappropriate content. The input is the feature vectors, and the output is a list of inappropriate scenes with timestamps and tags.

[0990] Step 6:

[0991] The device uses an emotion recognition engine to collect user emotional data. It uses cameras and sensors connected to the device to capture the user's facial expressions and heart rate in real time. The collected data is sent to a server. The input is raw data from the video and heart rate sensors, and the output is emotional data collected in real time.

[0992] Step 7:

[0993] The server analyzes the collected emotion data and associates it with inappropriate scenes. The emotion recognition engine analyzes the collected data and identifies emotional changes in specific scenes. For example, it identifies when a person's heart rate spikes and associates this with inappropriate scenes. The input is emotion data, and the output is emotion data associated with inappropriate scenes.

[0994] Step 8:

[0995] The server displays the timestamps and tags of inappropriate scenes, as well as emotion data, on a dashboard. It uses React.js to display the information in a web interface that parents can access. The input is inappropriate scenes and emotion data, and the output is the data displayed on the dashboard.

[0996] Step 9:

[0997] The parent logs in to the dashboard and checks the inappropriate scenes. After checking, the parent "approves" or "rejects" the scene correction. The input is the data displayed on the dashboard, and the output is the parent's operation data.

[0998] Step 10:

[0999] The server starts the editing module and converts the content of the inappropriate scene into an appropriate one. Specifically, it uses FFmpeg commands to process the scene that needs to be edited. The input is the parent operation data and the inappropriate scene information, and the output is the edited scene.

[1000] Step 11:

[1001] The server generates a new video file containing the modified content. A new edited video file is created and saved to storage. The input is the edited scene, and the output is the new video file.

[1002] Step 12:

[1003] The server provides the generated new video file to the parent's device. A download link is generated on the dashboard, which the parent can access to download the new video. The input is the new video file, and the output is the download link.

[1004] Step 13:

[1005] The parent plays the edited video on the device and confirms that it is safe for the child to watch. The input is the new video file, and the output is the parent's confirmation result.

[1006] Step 14:

[1007] The emotion recognition engine also collects emotional data when parents approve or reject corrections, and the server uses this data for future learning. This improves the accuracy of identifying inappropriate scenes from the next time onwards. The input is the parent's emotional data, and the output is learning data.

[1008] (Application example 2)

[1009] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1010] In conventional digital content viewing systems, it is difficult for users to manually detect and correct inappropriate scenes in the content they are viewing. Simply deleting inappropriate scenes often detracts from the user's viewing experience. Furthermore, there is a lack of a mechanism for more precisely identifying inappropriate scenes using user emotional data. This makes it difficult for users to enjoy content safely and comfortably.

[1011] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1012] In this invention, the server includes a means for a user to select or upload digital content they wish to view, a means for saving and analyzing digital content files, and a machine learning engine that extracts the analyzed digital content content as a feature vector and identifies scenes containing inappropriate content. This makes it possible to identify inappropriate scenes using emotion data. The server also includes a means for tagging inappropriate scenes identified by the machine learning engine and notifying a parental interface, a means for a parent to confirm the inappropriate scenes and approve corrections, and an editing module for converting the inappropriate scenes into appropriate content, thereby providing a safe and comfortable viewing experience for users. Furthermore, the server includes a means for collecting user emotion data and associating it with inappropriate scenes using the emotion engine, and a means for collecting emotion data when a parent approves or rejects corrections and performing learning based on this data, making it possible to more precisely identify inappropriate scenes and make appropriate corrections.

[1013] "Digital content" means media, including video, audio, images, and text, that is stored and distributed in electronic form.

[1014] A "machine learning engine" is a program that contains algorithms that analyze data, automatically learn patterns and features without relying on human-defined rules, and make predictions and classifications.

[1015] "Parental interface" refers to a user interface used by a parent to notify the parent of an inappropriate scene, and to review and approve the correction.

[1016] "Inappropriate content" refers to scenes that contain content that may have an upsetting or harmful effect on viewers, such as violence, discrimination, or extreme tragedy.

[1017] The "emotion engine" is a system that uses cameras and sensors to acquire biometric information such as a user's facial expressions and heart rate, and analyzes their emotional state.

[1018] An "editing module" is a software component for converting the content of identified inappropriate scenes into appropriate content.

[1019] A "feature vector" is data that quantifies the content extracted from the frames and audio data of digital content.

[1020] An "algorithm" is a set of steps or a computational method for solving a particular problem.

[1021] A "timestamp" is data that indicates the time at which a particular frame or scene occurred.

[1022] "Learning means" is the process by which a machine learning engine improves its model based on collected data.

[1023] MODE FOR CARRYING OUT THE INVENTION

[1024] To implement this invention, the following configuration and operation procedures are required.

[1025] The system of the present invention includes a terminal for users to select or upload digital content they wish to view, and a server for storing and analyzing digital content files. The server is equipped with a machine learning engine and an emotion engine, and collects and analyzes user emotion data in real time.

[1026] First, a user uploads digital content to a device. The device then sends the digital content file to a server. The server stores the received digital content file in temporary storage and initiates an analysis job. The server then transfers the digital content file to a machine learning engine, which extracts feature vectors from video frames and audio data. This allows a pre-trained algorithm to determine whether the content contains inappropriate language.

[1027] At the same time, the user's device is equipped with a camera and sensors, which allow the emotion engine to capture emotional data such as the user's facial expressions and heart rate. When an inappropriate scene is identified, the scene's timestamp and scene tag are recorded. When an inappropriate scene is identified, this information is notified to the parental interface.

[1028] Parents can review inappropriate scenes and approve their edits via the interface. Once the edits are approved, the server launches an editing module to convert the inappropriate content into appropriate content. For example, in a scene that contains a large amount of blood, the color of the blood can be changed to a harmless liquid, or the scene itself can be replaced with softer content.

[1029] Once edits are complete, the digital content is generated as a new file, which the server then provides to the user's device. The user can then play the edited digital content for safe and comfortable viewing. The emotion engine also collects emotional data from parents when they approve or reject the edits, which the machine learning engine can use to learn from to help identify inappropriate scenes in the future.

[1030] As a concrete example, consider the case where a children's animated content is uploaded. The system analyzes every frame and detects violent or tragic scenes. Parents can then review the content of the scenes on the dashboard and approve the edits. The edited video can then be safely shown to children.

[1031] An example of a prompt statement can be specified as follows:

[1032] Given the following text, generate a step-by-step guide to detecting and correcting inappropriate scenes in a children's cartoon, including collecting emotion data and analyzing the scene using AI.

[1033] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1034] Step 1:

[1035] The user selects or uploads digital content on the device.

[1036] Input: Digital content files (video files, audio files, etc.)

[1037] Output: The digital content file is saved to the server.

[1038] Specific operation: The user selects digital content through the terminal interface and clicks the upload button, after which the file is sent to the server.

[1039] Step 2:

[1040] The server stores the digital content file in temporary storage and starts the analysis job.

[1041] Input: Digital content file submitted by the user

[1042] Output: The digital content file is saved to temporary storage and analysis begins.

[1043] What happens: The server receives the digital content file and automatically stores it in temporary storage, then prepares it for analysis script execution.

[1044] Step 3:

[1045] The server transfers the digital content files to a machine learning engine, which extracts feature vectors for the video frames and audio data.

[1046] Input: Digital content file

[1047] Output: Feature vector (frame data, numerical representation of audio data)

[1048] Specific operation: The server transfers the digital content file to the machine learning engine, which processes the file to extract feature vectors from the frame data and audio data.

[1049] Step 4:

[1050] A machine learning engine uses feature vectors to detect scenes containing inappropriate language and record them with a timestamp.

[1051] Input: feature vector

[1052] Output: timestamp and tag data for the incorrect scene

[1053] How it works: The machine learning engine uses trained algorithms to analyze feature vectors and identify scenes that contain inappropriate content, such as violence or discrimination.

[1054] Step 5:

[1055] The device collects user emotional data in real time through cameras and sensors, and the emotion engine analyzes it.

[1056] Input: User's facial expression data, heart rate data

[1057] Output: Emotional data (stress response, heart rate variability, etc.)

[1058] Specific operation: The device's camera and sensors collect the user's facial expressions and biometric information, which are then sent to the emotion engine for analysis.

[1059] Step 6:

[1060] The emotion engine associates it with an inappropriate scene and sends it to the server.

[1061] Input: Emotion data, timestamps and tag data of inappropriate scenes

[1062] Output: Emotion data and association information for inappropriate scenes

[1063] Specific operation: The emotion engine matches the user's emotion data with the timestamp of the inappropriate scene based on the log.

[1064] Step 7:

[1065] The parental interface will be notified of inappropriate scenes and asked for approval to make corrections.

[1066] Input: timestamps and tag data of inappropriate scenes, emotion data

[1067] Output: Parental notification, request for approval of corrections

[1068] Specific behavior: The server displays detailed information about the inappropriate scene in the parental interface and asks the parent to approve or reject the correction.

[1069] Step 8:

[1070] After receiving parental approval, the server launches an editing module and converts the inappropriate scenes into appropriate content.

[1071] Input: Parental edit approval, timestamp and tag data for inappropriate scenes

[1072] Output: Edited scene data

[1073] Specific operation: The server's editing module converts the specified inappropriate scene into an appropriate expression and generates new scene data.

[1074] Step 9:

[1075] The edited digital content is generated as a new file and provided to the user's terminal.

[1076] Input: Edited scene data, original digital content excluding inappropriate scenes

[1077] Output: Edited digital content file

[1078] Specific operation: The server integrates the edited scene data into the entire original content, generates a new digital content file, and sends it to the user's device.

[1079] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1080] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1081] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1082] [Fourth embodiment]

[1083] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1084] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1085] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1086] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1087] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1088] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1089] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1090] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1091] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1092] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1093] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1094] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1095] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1096] The present invention relates to a system for providing an environment in which users can safely and appropriately watch videos they wish to watch. As shown below, the processing contents of a program in which the server, terminal, and user each play an important role are described.

[1097] First, a user logs in to the system's web application and selects or uploads the video they want to watch. At this time, the device sends the selected video file to the server. The server stores the received video file in temporary storage and starts the analysis job.

[1098] The server then transfers the video file to an AI engine for analysis. The AI ​​engine analyzes the video frames and audio data to extract feature vectors. This determines whether the video contains inappropriate content. For example, scenes in which a character attacks another character with a weapon or scenes containing large amounts of blood are identified as inappropriate content.

[1099] The server then displays the timestamps and scene tags of inappropriate scenes identified by the AI ​​engine on the parent's dashboard. Parents can log in to the dashboard and review the inappropriate scenes. If the parent approves the modification of the inappropriate scene, the server launches an editing module to convert the inappropriate content into an appropriate one. For example, in a scene depicting a large amount of blood, the color of the blood can be changed to a harmless liquid, or the scene itself can be replaced with softer content.

[1100] Once the editing is complete, the video is generated as a new file, which the server provides to the parent's device. The parent can then play the edited video on the device and confirm that the child is safe to watch.

[1101] As a concrete example, consider the case where a parent uploads a "children's cartoon" to the system. The system receives the video and identifies violent scenes as a result of analysis. The violent scenes are displayed on the dashboard, and when the parent clicks the approval button, the scenes are modified. For example, in a scene where a character slashes an enemy with a sword, the sword is changed to a soft stick and the depiction of the attack itself is deleted. This edited video is then provided to the parent's device, allowing the child to watch it safely.

[1102] This invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. In addition, the automatic analysis and editing functions using AI make it possible to provide an appropriate video environment without any hassle.

[1103] The processing flow will be explained below.

[1104] Step 1:

[1105] Users log into the system's web application and select or upload the video they want to watch.

[1106] Step 2:

[1107] The terminal transmits the selected or uploaded video file to the server.

[1108] Step 3:

[1109] The server stores the received video file in temporary storage.

[1110] Step 4:

[1111] The server adds a job to the queue to start video analysis.

[1112] Step 5:

[1113] The server runs the analysis job and transfers the video file to the artificial intelligence engine.

[1114] Step 6:

[1115] The artificial intelligence engine analyzes video frames and extracts feature vectors for each frame.

[1116] Step 7:

[1117] The artificial intelligence engine also analyzes voice data and converts it into text data.

[1118] Step 8:

[1119] The artificial intelligence engine uses pre-trained models to determine whether or not there is inappropriate language.

[1120] Step 9:

[1121] The artificial intelligence engine then timestamps and tags identified inappropriate scenes.

[1122] Step 10:

[1123] The server notifies the parent's dashboard of the tagged inappropriate scenes.

[1124] Step 11:

[1125] The user (parent) logs in to the dashboard and checks for inappropriate scenes.

[1126] Step 12:

[1127] The user (parent) approves or rejects the correction of the displayed inappropriate scene.

[1128] Step 13:

[1129] The server, upon receiving parental approval, queues a job to convert the inappropriate scenes to appropriate content.

[1130] Step 14:

[1131] The server launches the editing module to process the scene to be modified.

[1132] For example, frames depicting large amounts of blood could be converted to harmless versions or the scene could be removed entirely.

[1133] Step 15:

[1134] The server also edits the audio data to the appropriate content, maintaining overall synchronization.

[1135] Step 16:

[1136] The server generates a new video file after the editing is complete.

[1137] Step 17:

[1138] The server uploads the edited video to the user's dashboard and notifies them.

[1139] Step 18:

[1140] The user (parent) can check the edited video from the dashboard and play it on their device.

[1141] Step 19:

[1142] The user (child) can safely watch the edited video.

[1143] Example 1

[1144] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1145] There are many videos on the Internet that contain inappropriate content for children, making it difficult for parents to manage the content their children view. Editing out inappropriate scenes individually takes time and effort. To solve this problem, a system is needed that can automatically detect inappropriate content and edit it safely.

[1146] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1147] In this invention, the server includes a means for users to select or upload videos they wish to watch, a means for saving the video files in temporary storage and creating analysis jobs, and a means for transferring the video files to an artificial intelligence engine and requesting analysis. This makes it possible to automatically analyze videos containing inappropriate content and edit them in a safe manner.

[1148] "User" means an individual or entity that uses the System to select or upload videos that they wish to view.

[1149] A "video file" is a digital data file that contains visual and audio content.

[1150] "Temporary storage" is a digital storage device for temporarily storing video files until an analysis job is started.

[1151] An "analysis job" is a processing task that transfers a video file to an artificial intelligence engine and has its contents analyzed.

[1152] An "artificial intelligence engine" is a software or hardware system that analyzes video frames and audio data and runs algorithms to identify inappropriate language.

[1153] A "feature vector" is data that expresses each feature extracted from video frames or audio as a numerical value.

[1154] "Inappropriate content" refers to video scenes that contain violence, blood, or obscene content that is undesirable for children or general audiences.

[1155] A "timestamp" is data indicating the occurrence time of a frame containing an inappropriate expression.

[1156] A "scene tag" is a label that is assigned to identify information related to a particular scene.

[1157] "Parents" are guardians who are responsible for managing the content their children view.

[1158] The "overview panel" is an interface provided by the system, which is a dashboard for users to check the video analysis results and inappropriate scenes.

[1159] An "editing module" is a software or hardware system for changing detected inappropriate scenes into appropriate content.

[1160] "User Terminal" means a device used by a parent or user to access the system and view or manage videos.

[1161] A "new file" is a digital file generated containing the video modified by the editing module.

[1162] The present invention provides a system for providing an environment in which users can safely and appropriately watch videos they wish to watch. This system is realized using the following specific hardware and software.

[1163] 1. User video selection and upload

[1164] Users log in to the system's web application and select or upload the video they want to watch. To upload, the device sends the video file to the server using an HTTP POST request.

[1165] 2. Video transmission and storage to the server

[1166] The video file sent by the device is received by the server and saved in temporary storage (e.g., AWS S3 bucket, Google Cloud Storage), ensuring safe storage until the analysis job begins.

[1167] 3. Start video analysis

[1168] The server obtains the path of the saved video file, creates an analysis job, and adds it to the queue. A message broker such as RabbitMQ or Apache Kafka is used for the analysis job queue. The server then transfers the analysis job to the AI ​​engine and requests it to be analyzed.

[1169] 4. Video analysis using an AI engine

[1170] The AI ​​engine splits the video file into frames and analyzes the frames and audio data in parallel using deep learning models (e.g., YOLO, ResNet, and WaveNet for audio analysis). It extracts feature vectors from each frame and audio and determines whether the features represent inappropriate content (e.g., violence, blood, obscenity).

[1171] 5. Dashboard display of inappropriate scene information

[1172] The server generates data to display the analysis results from the AI ​​engine on the parent's dashboard. When the parent logs in, this data is displayed and the parent can check the timestamps of inappropriate scenes and scene tags.

[1173] 6. Parental approval of inappropriate scenes and corrections

[1174] Parents can check inappropriate scenes via the dashboard and approve corrections by sending a correction request to the server.

[1175] 7. Edit video with the editing module

[1176] After receiving parental approval, the server launches the editing module, which applies filters and replacements to inappropriate scenes based on the AI ​​engine's judgment. These modifications include changing the color of blood and muting obscene content.

[1177] 8. Video generation and delivery to devices

[1178] The modified video file is created as a new file and saved in temporary storage. The server sends a link to this file to the parent device and notifies them that editing is complete.

[1179] Specific examples

[1180] When uploading a "children's cartoon," the system analyzes the video and identifies violent scenes. If the parent reviews the details of the violent scene displayed on the dashboard and approves the edits, a scene in which a character slashes an enemy with a sword is changed to a soft stick, and the attack is deleted. A new file is generated and provided to the parent's device. An example prompt is, "Please explain the system for editing and providing children's videos to make them safe."

[1181] This invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. In addition, the automatic analysis and editing functions using AI make it possible to provide an appropriate video environment without any hassle.

[1182] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1183] Step 1:

[1184] User selects or uploads a video

[1185] A user logs in to a web application and selects or uploads a video they want to watch. When the user clicks the upload button for a video file, the device sends the video file to the server using an HTTP POST request. The input is the video file selected by the user, and the output is the transmission of the video file to the server. Specifically, the device reads the metadata of the video file, generates an upload request, and sends it to the server.

[1186] Step 2:

[1187] Sending and saving video files to the server

[1188] The video file sent from the device is received by the server. The server saves the received video file in temporary storage (e.g., AWS S3 bucket, Google Cloud Storage). The input is the video file sent from the device, and the output is the video file saved in temporary storage. Specifically, the server generates a path to save the video file and uploads the file to storage.

[1189] Step 3:

[1190] Starting a video analysis job

[1191] The server obtains the path of the saved video file, creates an analysis job, and adds it to a message queue. The input is the path of the video file saved in temporary storage, and the output is the job added to the analysis job queue. Specifically, the server generates a job ID and adds the job details, along with the path of the video file, to a queue (such as RabbitMQ or Apache Kafka).

[1192] Step 4:

[1193] Video analysis by AI engine

[1194] The server transfers the analysis job to the AI ​​engine and requests analysis. The AI ​​engine divides the video file into frames and analyzes the frame and audio data. The input is the analysis job details (video file path, job ID), and the output is the analysis result data (feature vector, timestamp, scene tag). Specifically, the AI ​​engine analyzes the images and audio using YOLO, ResNet, or WaveNet models, and extracts feature vectors from each frame and audio.

[1195] Step 5:

[1196] Dashboard display of inappropriate scene information

[1197] The server converts the analysis results received from the AI ​​engine into dashboard data and displays it on the parent dashboard. The input is the analysis result data of inappropriate scenes, and the output is the timestamp and scene tag displayed on the parent dashboard. Specifically, the server formats the data and sends it to the dashboard frontend using WebSocket or HTTP requests.

[1198] Step 6:

[1199] Parental review and approval of inappropriate scenes

[1200] Parents log in to the dashboard, check the inappropriate scenes, and approve their corrections. The input is a list of inappropriate scenes displayed on the dashboard, and the output is a request for approval of the corrections. Specifically, when a parent checks the scene details and thumbnails and clicks the "Approve Corrections" button, a request is sent to the server.

[1201] Step 7:

[1202] Editing module for video editing

[1203] The server launches the editing module after receiving approval for the correction from the parent. The editing module corrects the inappropriate scenes based on the AI ​​engine's judgment. The input is the correction approval request and details of the inappropriate scenes, and the output is the corrected video file. Specifically, the editing module detects inappropriate frames and applies filters and replacement processes to correct the content of the scenes.

[1204] Step 8:

[1205] Generate edited video and deliver it to the device

[1206] The server generates a new video file modified by the editing module and stores it in temporary storage. The input is the modified video data, and the output is a link to the new video file. Specifically, the server generates the video file and sends its path to the parent device.

[1207] Step 9:

[1208] Play and check the edited video

[1209] The parent uses the provided link to play the edited video on their device and confirm the content. The input is a link to the edited video file, and the output is the confirmed video playback. Specifically, the parent clicks the link to launch the video player and watches the edited video to confirm that it is safe for their child to watch.

[1210] (Application example 1)

[1211] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1212] With the recent spread of video streaming services, the safety of children's content has become a major issue. In particular, it is important to prevent videos containing inappropriate content from falling into the hands of children. However, it takes a great deal of time and effort for parents to manually review and edit the content of all videos. To solve this problem, an efficient and highly accurate automated system is needed.

[1213] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1214] In this invention, the server includes: means for a user to select or upload a video they wish to watch; a server for storing and analyzing video files; an artificial intelligence engine for extracting the analyzed video content as a feature vector and identifying scenes containing inappropriate content; means for tagging the inappropriate scenes identified by the artificial intelligence engine and notifying the user on their dashboard; means for the user to review the inappropriate scenes and approve their correction; an editing module for converting the inappropriate scenes into appropriate content; and means for generating edited videos and providing them to a content distribution service. This allows users to quickly and easily manage, edit, and provide safe videos for children without any hassle.

[1215] "Means for users to select or upload videos they wish to view" refers to a function that allows users to use their own devices to specify or upload videos they wish to view to the system.

[1216] A "server for storing and analyzing video files" is a computer device that temporarily stores uploaded video files and analyzes their contents using an artificial intelligence engine.

[1217] The "artificial intelligence engine for extracting the analyzed video content as feature vectors and identifying scenes containing inappropriate expressions" is a processing device equipped with an artificial intelligence algorithm for analyzing video frames and audio data and automatically detecting inappropriate scenes based on the resulting feature vectors.

[1218] "Means for tagging inappropriate scenes identified by the AI ​​engine and notifying the user on their dashboard" refers to a mechanism for tagging inappropriate scenes detected by the AI ​​engine and displaying that information on a dashboard where users can view it.

[1219] "Means for users to confirm inappropriate scenes and approve corrections" refers to a function that allows users to use the dashboard to confirm the content of inappropriate scenes and select and approve whether or not to correct the scenes.

[1220] The "editing module for converting inappropriate scenes into appropriate content" is a software module that has an editing function for converting scenes that are determined to be inappropriate into more appropriate content based on the user's approval.

[1221] "Means for generating edited video and providing it to a content distribution service" refers to a function for generating a new video file containing the modified scene and providing that video to users via a content distribution service.

[1222] The present invention relates to a system that provides a safe viewing environment for video content for children. In this system, the server, the terminal, and the user each play an important role.

[1223] First, a user uploads or selects a video they wish to watch using their device. The video file is then sent to the server and temporarily stored in storage. The server then transfers the received video file to an AI engine for analysis. The AI ​​engine analyzes the video frames and audio data to extract feature vectors. This allows it to identify scenes that contain inappropriate content, such as violent scenes or scenes containing inappropriate language.

[1224] The analysis results are extracted as timestamps and scene tags for inappropriate scenes identified by the AI ​​engine. This information is then posted to the user's dashboard. Parents or appropriate administrators can log in to the dashboard to review the inappropriate scenes and approve their corrections. For example, if a character uses a sword in a violent scene, the scene can be changed to a soft stick, or the scene itself can be replaced with different content.

[1225] Once approval is given, an editing module is launched on the server side and the inappropriate scenes are converted into appropriate content. This editing process is carried out automatically using artificial intelligence. The edited video is generated as a new file and provided to the user's device via the content distribution service. This ensures that children can watch videos safely.

[1226] The following hardware and software can be used to build the system. It is recommended that the server be equipped with storage, a CPU, and a GPU for storing and analyzing video files. Deep learning frameworks such as TensorFlow and PyTorch can be used as the artificial intelligence engine. Terminals include devices such as smartphones, tablets, and PCs.

[1227] As a concrete example, consider the case where a user uploads a children's cartoon. The system receives the video and identifies violent scenes as a result of analysis by an artificial intelligence engine. The dashboard displays timestamps and scene tags for violent scenes, and parents can click a correction button to edit the scene. For example, in a scene where a character slashes an enemy with a sword, the sword can be changed to a soft stick. In this way, an edited video is generated and provided to the user's device via a distribution service.

[1228] Examples of prompts include:

[1229] In order to analyze videos for inappropriate content and provide a safe viewing environment, please analyze and edit your videos according to the following requirements.

[1230] Identify any inappropriate scenes and provide their start and end times

[1231] If a scene needs editing, replace it with appropriate content.

[1232] In this way, users can quickly and easily manage, edit and provide safe videos for children without any hassle.

[1233] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1234] Step 1:

[1235] A user uses a terminal to upload the video they want to watch to the system. The input is the video file selected or uploaded by the user, and the output is the video file sent to the server, at which point the file is saved in temporary storage.

[1236] Step 2:

[1237] The server transfers the received video file to the AI ​​engine and begins analysis. The input is the video file stored on the server, and the output is the video data sent to the AI ​​engine. The server completes the video file transfer procedure and prepares for analysis.

[1238] Step 3:

[1239] The AI ​​engine analyzes video frames and audio data to extract feature vectors. The input is the video data transferred from the server, and the output is the analysis result: feature vectors, timestamps of inappropriate scenes, and scene tags. An AI model (e.g., TensorFlow or PyTorch) is used to analyze each frame in the video and perform data calculations to detect inappropriate content.

[1240] Step 4:

[1241] The server tags inappropriate scenes identified by the AI ​​engine and notifies the user's dashboard. The input is the analysis result from the AI ​​engine, and the output is the inappropriate scene information notified on the dashboard. The server reflects the tag information and timestamp on the dashboard.

[1242] Step 5:

[1243] The user logs in to the dashboard, checks the inappropriate scenes, and approves the corrections. The input is the inappropriate scene information displayed on the dashboard, and the output is a command to approve the corrections. The user's operation determines whether the inappropriate scenes need to be corrected.

[1244] Step 6:

[1245] The server starts the editing module based on the approved correction request. The input is a command from the user to approve the correction, and the output is the start of the editing module. The editing module automatically converts inappropriate scenes into appropriate ones.

[1246] Step 7:

[1247] The editing module converts inappropriate scenes into appropriate content and generates a new edited video file. The input is the inappropriate scene data handled by the editing module, and the output is the edited video file. The AI ​​model is used to correct the content of the scene and generate a new edited video.

[1248] Step 8:

[1249] The server saves the edited video file as a final file and provides it to the user's device through the content distribution service. The input is the edited video file, and the output is the video file provided to the user's device through the distribution service. Users can receive safe video content for children to watch.

[1250] This allows users to easily create a safe video viewing environment for children.

[1251] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1252] The present invention relates to a system for providing an environment in which users can watch videos they want to watch in a safe and appropriate manner. In particular, the present invention describes a system that combines an emotion engine that recognizes the user's emotions and efficiently identifies and corrects inappropriate scenes.

[1253] The system begins with the user selecting or uploading a video they want to watch. Once the user uploads the video, the device sends the video file to the server. The server then stores the received video file in temporary storage and starts the analysis job.

[1254] The server transfers the video file to an AI engine, which extracts feature vectors from the video frames and audio data. This allows a pre-trained model to determine whether the video contains inappropriate content. For example, scenes in which a character attacks another character with a weapon or contains a large amount of blood are identified as inappropriate content.

[1255] Furthermore, the present invention incorporates an emotion engine that collects emotional data from the user while watching. The emotion engine acquires emotional data such as the user's facial expressions and heart rate through cameras and sensors, and associates the data with inappropriate scenes. For example, if the user's heart rate temporarily rises or if a stressed reaction is observed in the user's facial expression, the emotion engine determines that the scene is likely to be inappropriate.

[1256] The server displays the timestamps and scene tags of inappropriate scenes identified by the AI ​​engine, as well as emotional data from the emotion engine, on the parent's dashboard. Parents log in to the dashboard and review the inappropriate scenes. If the parent approves the correction of a confirmed scene, the server launches an editing module to convert the inappropriate content into something appropriate. For example, in a scene depicting a large amount of blood, the color of the blood can be changed to a harmless liquid, or the scene itself can be replaced with softer content.

[1257] The edited video is then generated as a new file, which the server then provides to the parent's device. The parent then plays the edited video to ensure it is safe for their child to watch. The emotion engine also collects emotional data when the parent approves or rejects the edits, allowing the AI ​​engine to learn from this data to help identify inappropriate scenes in the future.

[1258] As a concrete example, consider the case where a parent uploads a "family movie" to the system. The system receives the video and analyzes it using AI and an emotion engine. The emotion engine collects user emotional data and determines violent or tragic scenes as inappropriate. Based on this information, the parent approves the scene modification on the dashboard, and the AI ​​engine converts the scene into appropriate content, resulting in a safe video.

[1259] In this way, this invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. Furthermore, by combining AI and an emotion engine, it is possible to identify inappropriate scenes with greater accuracy and make appropriate corrections.

[1260] The processing flow will be explained below.

[1261] Step 1:

[1262] Users log into the system's web application and select or upload the video they want to watch.

[1263] Step 2:

[1264] The terminal transmits the selected or uploaded video file to the server.

[1265] Step 3:

[1266] The server stores the received video file in temporary storage.

[1267] Step 4:

[1268] The server adds a job to the queue to start video analysis.

[1269] Step 5:

[1270] The server runs the analysis job and transfers the video file to the artificial intelligence engine.

[1271] Step 6:

[1272] The artificial intelligence engine analyzes video frames and extracts feature vectors for each frame.

[1273] Step 7:

[1274] The artificial intelligence engine analyzes the voice data and converts it into text data.

[1275] Step 8:

[1276] The artificial intelligence engine uses pre-trained models to determine whether a sentence contains inappropriate language.

[1277] Step 9:

[1278] The artificial intelligence engine will timestamp and tag inappropriate scenes.

[1279] Step 10:

[1280] The emotion engine collects emotional data while users watch videos, for example by analyzing their facial expressions using a camera and monitoring their heart rate with a heart rate sensor.

[1281] Step 11:

[1282] The emotion engine analyzes the collected emotion data and associates it with inappropriate scenes.

[1283] Step 12:

[1284] The server notifies the parent's dashboard of tagged inappropriate scenes and associated emotional data.

[1285] Step 13:

[1286] The user (parent) logs in to the dashboard and checks the inappropriate scenes and their emotion data.

[1287] Step 14:

[1288] The user (parent) approves or rejects the correction of the inappropriate scene.

[1289] Step 15:

[1290] The server, upon receiving parental approval, queues a job to convert the inappropriate scenes to appropriate content.

[1291] Step 16:

[1292] The server launches an editing module to process the scene to be altered, for example, converting a frame depicting large amounts of blood into a harmless liquid and replacing the scene with a different musical melody.

[1293] Step 17:

[1294] The server also edits the audio data to the appropriate content, maintaining overall synchronization.

[1295] Step 18:

[1296] The server generates a new video file after the editing is complete.

[1297] Step 19:

[1298] The server uploads the edited video to the user's dashboard and notifies them.

[1299] Step 20:

[1300] The user (parent) can check the edited video from the dashboard and play it on their device.

[1301] Step 21:

[1302] The user (child) can safely watch the edited video.

[1303] Step 22:

[1304] The emotion engine also collects emotional data when parents approve or reject modifications and feeds this back to the AI ​​engine.

[1305] Step 23:

[1306] The artificial intelligence engine uses the collected emotional data to optimize algorithms for identifying inappropriate scenes.

[1307] Example 2

[1308] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1309] Currently, when a scene containing inappropriate content is found in video content viewed by children, parents must manually detect and edit that scene. This places a significant burden on parents, and there are limitations to the accuracy of identifying and editing inappropriate scenes. Furthermore, there is no established method for effectively utilizing user emotion data to improve the accuracy of identifying inappropriate scenes. Therefore, there is a need for a system that provides a safe and appropriate video viewing environment and reduces the burden on parents.

[1310] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for selecting or uploading a video that a user wants to watch; means for saving and analyzing video files; means for extracting the content of the analyzed video as a feature vector and identifying scenes containing inappropriate expressions; means for collecting user emotion data and associating it with inappropriate scenes; means for tagging inappropriate scenes identified by the artificial intelligence engine and the emotion recognition engine and notifying them to a parent's dashboard; means for a parent to confirm and approve corrections of inappropriate scenes; means for converting inappropriate scenes into appropriate content; and means for generating edited videos and providing them to a user terminal. This enables automatic detection and editing of scenes containing inappropriate expressions, significantly reducing the burden on parents and providing an environment in which children can watch safe and appropriate videos.

[1311] "User" refers to the person who selects or uploads videos to the system that they wish to view.

[1312] "Video File" means a digital file containing video and audio data selected or uploaded by a User for viewing purposes.

[1313] A "server" refers to a computer system that stores video files and handles a series of processes such as analyzing and editing them.

[1314] "Feature vector" refers to vector data that quantifies specific attributes and features extracted from video frames and audio.

[1315] "Artificial intelligence engine" refers to a machine learning model that uses feature vectors to analyze videos and identify scenes containing inappropriate content.

[1316] An "emotion recognition engine" refers to a system that uses cameras and sensors to collect user emotional data and associates it with inappropriate scenes.

[1317] "Dashboard" refers to the user interface that allows parents to access the system to review information about inappropriate scenes and approve corrections.

[1318] "Editing Module" refers to software functionality for converting identified inappropriate scenes into appropriate content.

[1319] A "timestamp" refers to data that indicates the specific time at which a particular frame or scene appears in a video.

[1320] "Tags" refer to metadata used to classify and identify inappropriate scenes.

[1321] The present invention relates to a system for providing an environment in which a user can safely and appropriately view videos that the user desires to view. Detailed embodiments of the present invention will be described below.

[1322] First, the user selects or uploads the video they want to watch. The device sends the video file uploaded by the user to the server. The device can be a general personal computer or smartphone, and uses a file selection dialog that runs on a web browser.

[1323] Next, the server temporarily stores the received video file in a cloud storage service such as Amazon S3.

[1324] The server queues an analysis job for the stored video file and starts executing the job. The server sends the video file to an artificial intelligence engine (e.g., a TensorFlow model) that extracts feature vectors from the video frames and audio data.

[1325] The AI ​​engine uses feature vectors, specifically pre-trained machine learning models, to identify inappropriate scenes, such as scenes of violence or excessive gore.

[1326] Furthermore, the system collects user emotional data using cameras and sensors connected to the device. A standard webcam (e.g., Logitech C920) is used as the camera, and a heart rate monitor (e.g., Fitbit Charge 4) is used as the sensor. The Affectiva SDK is used as the emotion recognition engine. This allows the system to capture the user's facial expressions and heart rate in real time and associate them with inappropriate scenes.

[1327] The server displays the collected emotion data and timestamps and tags of inappropriate scenes on a dashboard. The dashboard is implemented using a web framework such as React.js. Parents can log in to the dashboard and check the inappropriate scenes.

[1328] If a parent approves the correction of an inappropriate scene, the server launches an editing module (e.g., FFmpeg command) that converts the inappropriate content into something appropriate, such as changing the color of blood or cutting the scene altogether.

[1329] Once the editing is complete, the video is generated as a new file, which is then provided to the device by the server. A link is displayed on the dashboard for parents to download the new video file. When parents click the link, the file is downloaded to the device.

[1330] Finally, parents can play the edited video on their device to confirm that it is safe for their child to watch. The emotion recognition engine also collects emotional data when parents approve or reject the edits, and uses this data for future learning on the server. This improves the accuracy of identifying inappropriate scenes in future videos.

[1331] As a concrete example, consider the case where a parent uploads a "family movie" to the system. The system receives the video and analyzes it using AI and an emotion recognition engine. For example, it uses a prompt such as, "Please detect whether this movie contains violent or tragic scenes." The emotion recognition engine collects the user's emotional data and determines that violent or tragic scenes are inappropriate. Based on this information, if the parent approves the scene modification on the dashboard, the AI ​​engine converts the scene into appropriate content, and a safe video is generated.

[1332] In this way, this invention allows parents to easily manage the content of videos their children watch and avoid scenes containing inappropriate content. Furthermore, by combining AI and an emotion recognition engine, it is possible to identify inappropriate scenes with greater accuracy and make appropriate corrections.

[1333] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1334] Step 1:

[1335] The user selects or uploads the video they want to watch. The device retrieves the video file selected or uploaded by the user and sends it to the server using an HTTP POST request. The input is the video file selected by the user, and the output is the video file sent to the server.

[1336] Step 2:

[1337] The server saves the received video file in temporary storage. Specifically, it uploads the video file to cloud storage such as Amazon S3 and records the status of the saving completion in a log. The input is the video file sent to the server, and the output is the video file saved in the storage.

[1338] Step 3:

[1339] The server queues an analysis job for the stored video file and starts the job execution. The video file is sent to the artificial intelligence engine for analysis. The input is the video file stored in storage, and the output is the start status of the analysis job.

[1340] Step 4:

[1341] The server transfers the video file to an AI engine, which extracts feature vectors from the video frames and audio data. The AI ​​engine (e.g., a TensorFlow model) converts each video frame and audio data into a numerical vector and inputs it into a generative AI model. The input is the video file and its frames and audio data, and the output is the feature vector.

[1342] Step 5:

[1343] The AI ​​engine uses feature vectors to identify inappropriate scenes. Specifically, it applies a pre-trained model and uses a specific algorithm to detect frames containing inappropriate content. The input is the feature vectors, and the output is a list of inappropriate scenes with timestamps and tags.

[1344] Step 6:

[1345] The device uses an emotion recognition engine to collect user emotional data. It uses cameras and sensors connected to the device to capture the user's facial expressions and heart rate in real time. The collected data is sent to a server. The input is raw data from the video and heart rate sensors, and the output is emotional data collected in real time.

[1346] Step 7:

[1347] The server analyzes the collected emotion data and associates it with inappropriate scenes. The emotion recognition engine analyzes the collected data and identifies emotional changes in specific scenes. For example, it identifies when a person's heart rate spikes and associates this with inappropriate scenes. The input is emotion data, and the output is emotion data associated with inappropriate scenes.

[1348] Step 8:

[1349] The server displays the timestamps and tags of inappropriate scenes, as well as emotion data, on a dashboard. It uses React.js to display the information in a web interface that parents can access. The input is inappropriate scenes and emotion data, and the output is the data displayed on the dashboard.

[1350] Step 9:

[1351] The parent logs in to the dashboard and checks the inappropriate scenes. After checking, the parent "approves" or "rejects" the scene correction. The input is the data displayed on the dashboard, and the output is the parent's operation data.

[1352] Step 10:

[1353] The server starts the editing module and converts the content of the inappropriate scene into an appropriate one. Specifically, it uses FFmpeg commands to process the scene that needs to be edited. The input is the parent operation data and the inappropriate scene information, and the output is the edited scene.

[1354] Step 11:

[1355] The server generates a new video file containing the modified content. A new edited video file is created and saved to storage. The input is the edited scene, and the output is the new video file.

[1356] Step 12:

[1357] The server provides the generated new video file to the parent's device. A download link is generated on the dashboard, which the parent can access to download the new video. The input is the new video file, and the output is the download link.

[1358] Step 13:

[1359] The parent plays the edited video on the device and confirms that it is safe for the child to watch. The input is the new video file, and the output is the parent's confirmation result.

[1360] Step 14:

[1361] The emotion recognition engine also collects emotional data when parents approve or reject corrections, and the server uses this data for future learning. This improves the accuracy of identifying inappropriate scenes from the next time onwards. The input is the parent's emotional data, and the output is learning data.

[1362] (Application example 2)

[1363] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1364] In conventional digital content viewing systems, it is difficult for users to manually detect and correct inappropriate scenes in the content they are viewing. Simply deleting inappropriate scenes often detracts from the user's viewing experience. Furthermore, there is a lack of a mechanism for more precisely identifying inappropriate scenes using user emotional data. This makes it difficult for users to enjoy content safely and comfortably.

[1365] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1366] In this invention, the server includes a means for a user to select or upload digital content they wish to view, a means for saving and analyzing digital content files, and a machine learning engine that extracts the analyzed digital content content as a feature vector and identifies scenes containing inappropriate content. This makes it possible to identify inappropriate scenes using emotion data. The server also includes a means for tagging inappropriate scenes identified by the machine learning engine and notifying a parental interface, a means for a parent to confirm the inappropriate scenes and approve corrections, and an editing module for converting the inappropriate scenes into appropriate content, thereby providing a safe and comfortable viewing experience for users. Furthermore, the server includes a means for collecting user emotion data and associating it with inappropriate scenes using the emotion engine, and a means for collecting emotion data when a parent approves or rejects corrections and performing learning based on this data, making it possible to more precisely identify inappropriate scenes and make appropriate corrections.

[1367] "Digital content" means media, including video, audio, images, and text, that is stored and distributed in electronic form.

[1368] A "machine learning engine" is a program that contains algorithms that analyze data, automatically learn patterns and features without relying on human-defined rules, and make predictions and classifications.

[1369] "Parental interface" refers to a user interface used by a parent to notify the parent of an inappropriate scene, and to review and approve the correction.

[1370] "Inappropriate content" refers to scenes that contain content that may have an upsetting or harmful effect on viewers, such as violence, discrimination, or extreme tragedy.

[1371] The "emotion engine" is a system that uses cameras and sensors to acquire biometric information such as a user's facial expressions and heart rate, and analyzes their emotional state.

[1372] An "editing module" is a software component for converting the content of identified inappropriate scenes into appropriate content.

[1373] A "feature vector" is data that quantifies the content extracted from the frames and audio data of digital content.

[1374] An "algorithm" is a set of steps or a computational method for solving a particular problem.

[1375] A "timestamp" is data that indicates the time at which a particular frame or scene occurred.

[1376] "Learning means" is the process by which a machine learning engine improves its model based on collected data.

[1377] MODE FOR CARRYING OUT THE INVENTION

[1378] To implement this invention, the following configuration and operation procedures are required.

[1379] The system of the present invention includes a terminal for users to select or upload digital content they wish to view, and a server for storing and analyzing digital content files. The server is equipped with a machine learning engine and an emotion engine, and collects and analyzes user emotion data in real time.

[1380] First, a user uploads digital content to a device. The device then sends the digital content file to a server. The server stores the received digital content file in temporary storage and initiates an analysis job. The server then transfers the digital content file to a machine learning engine, which extracts feature vectors from video frames and audio data. This allows a pre-trained algorithm to determine whether the content contains inappropriate language.

[1381] At the same time, the user's device is equipped with a camera and sensors, which allow the emotion engine to capture emotional data such as the user's facial expressions and heart rate. When an inappropriate scene is identified, the scene's timestamp and scene tag are recorded. When an inappropriate scene is identified, this information is notified to the parental interface.

[1382] Parents can review inappropriate scenes and approve their edits via the interface. Once the edits are approved, the server launches an editing module to convert the inappropriate content into appropriate content. For example, in a scene that contains a large amount of blood, the color of the blood can be changed to a harmless liquid, or the scene itself can be replaced with softer content.

[1383] Once edits are complete, the digital content is generated as a new file, which the server then provides to the user's device. The user can then play the edited digital content for safe and comfortable viewing. The emotion engine also collects emotional data from parents when they approve or reject the edits, which the machine learning engine can use to learn from to help identify inappropriate scenes in the future.

[1384] As a concrete example, consider the case where a children's animated content is uploaded. The system analyzes every frame and detects violent or tragic scenes. Parents can then review the content of the scenes on the dashboard and approve the edits. The edited video can then be safely shown to children.

[1385] An example of a prompt statement can be specified as follows:

[1386] Given the following text, generate a step-by-step guide to detecting and correcting inappropriate scenes in a children's cartoon, including collecting emotion data and analyzing the scene using AI.

[1387] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1388] Step 1:

[1389] The user selects or uploads digital content on the device.

[1390] Input: Digital content files (video files, audio files, etc.)

[1391] Output: The digital content file is saved to the server.

[1392] Specific operation: The user selects digital content through the terminal interface and clicks the upload button, after which the file is sent to the server.

[1393] Step 2:

[1394] The server stores the digital content file in temporary storage and starts the analysis job.

[1395] Input: Digital content file submitted by the user

[1396] Output: The digital content file is saved to temporary storage and analysis begins.

[1397] What happens: The server receives the digital content file and automatically stores it in temporary storage, then prepares it for analysis script execution.

[1398] Step 3:

[1399] The server transfers the digital content files to a machine learning engine, which extracts feature vectors for the video frames and audio data.

[1400] Input: Digital content file

[1401] Output: Feature vector (frame data, numerical representation of audio data)

[1402] Specific operation: The server transfers the digital content file to the machine learning engine, which processes the file to extract feature vectors from the frame data and audio data.

[1403] Step 4:

[1404] A machine learning engine uses feature vectors to detect scenes containing inappropriate language and record them with a timestamp.

[1405] Input: feature vector

[1406] Output: timestamp and tag data for the incorrect scene

[1407] How it works: The machine learning engine uses trained algorithms to analyze feature vectors and identify scenes that contain inappropriate content, such as violence or discrimination.

[1408] Step 5:

[1409] The device collects user emotional data in real time through cameras and sensors, and the emotion engine analyzes it.

[1410] Input: User's facial expression data, heart rate data

[1411] Output: Emotional data (stress response, heart rate variability, etc.)

[1412] Specific operation: The device's camera and sensors collect the user's facial expressions and biometric information, which are then sent to the emotion engine for analysis.

[1413] Step 6:

[1414] The emotion engine associates it with an inappropriate scene and sends it to the server.

[1415] Input: Emotion data, timestamps and tag data of inappropriate scenes

[1416] Output: Emotion data and association information for inappropriate scenes

[1417] Specific operation: The emotion engine matches the user's emotion data with the timestamp of the inappropriate scene based on the log.

[1418] Step 7:

[1419] The parental interface will be notified of inappropriate scenes and asked for approval to make corrections.

[1420] Input: timestamps and tag data of inappropriate scenes, emotion data

[1421] Output: Parental notification, request for approval of corrections

[1422] Specific behavior: The server displays detailed information about the inappropriate scene in the parental interface and asks the parent to approve or reject the correction.

[1423] Step 8:

[1424] After receiving parental approval, the server launches an editing module and converts the inappropriate scenes into appropriate content.

[1425] Input: Parental edit approval, timestamp and tag data for inappropriate scenes

[1426] Output: Edited scene data

[1427] Specific operation: The server's editing module converts the specified inappropriate scene into an appropriate expression and generates new scene data.

[1428] Step 9:

[1429] The edited digital content is generated as a new file and provided to the user's terminal.

[1430] Input: Edited scene data, original digital content excluding inappropriate scenes

[1431] Output: Edited digital content file

[1432] Specific operation: The server integrates the edited scene data into the entire original content, generates a new digital content file, and sends it to the user's device.

[1433] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1434] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1435] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1436] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1437] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1438] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1439] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1440] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1441] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1442] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1443] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1444] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1445] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1446] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1447] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1448] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1449] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1450] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1451] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1452] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1453] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1454] The following is further disclosed regarding the above embodiment.

[1455] (Claim 1)

[1456] a means for a user to select or upload a video they wish to view;

[1457] a server for storing and analyzing video files;

[1458] An artificial intelligence engine that extracts the analyzed video content as a feature vector and identifies scenes containing inappropriate content;

[1459] A means to tag inappropriate scenes identified by the AI ​​engine and notify them on the parent dashboard;

[1460] A means for parents to review and approve corrections of inappropriate scenes;

[1461] an editing module for converting inappropriate scenes into appropriate content;

[1462] A means for generating an edited video and providing it to a user terminal;

[1463] A system including:

[1464] (Claim 2)

[1465] 2. The system of claim 1, further comprising means for the artificial intelligence engine to analyze video frames and audio data and detect inappropriate language using a specific algorithm.

[1466] (Claim 3)

[1467] 10. The system of claim 1, further comprising means for recording a timestamp and a scene tag of a frame containing profanity.

[1468] "Example 1"

[1469] (Claim 1)

[1470] a means for a user to select or upload a video they wish to view;

[1471] A server for storing video files in temporary storage and creating analysis jobs;

[1472] A means to transfer video files to an AI engine and request analysis,

[1473] An artificial intelligence engine that extracts the analyzed video content as a feature vector and identifies scenes containing inappropriate content;

[1474] a means for tagging the inappropriate scenes identified by the artificial intelligence engine with timestamps and scene tags and notifying them to a parent overview panel;

[1475] A means for parents to review and approve corrections of inappropriate scenes;

[1476] a means for activating an editing module and converting the inappropriate scenes into appropriate content;

[1477] A means for generating the edited video as a new file and providing it to a user terminal;

[1478] A system including:

[1479] (Claim 2)

[1480] 2. The system of claim 1, further comprising means for the artificial intelligence engine to analyze video frames and audio data and detect inappropriate language using a specific algorithm.

[1481] (Claim 3)

[1482] 10. The system of claim 1, further comprising means for recording timestamps and scene tags of frames containing profanity for display in the parent overview panel.

[1483] "Application Example 1"

[1484] (Claim 1)

[1485] a means for a user to select or upload a video they wish to view;

[1486] a server for storing and analyzing video files;

[1487] An artificial intelligence engine that extracts the analyzed video content as a feature vector and identifies scenes containing inappropriate content;

[1488] The AI ​​engine will tag inappropriate scenes identified and notify users on their dashboard.

[1489] A means for users to review and approve corrections of inappropriate scenes;

[1490] an editing module for converting inappropriate scenes into appropriate content;

[1491] a means for generating and providing the edited video to a content distribution service;

[1492] A system including:

[1493] (Claim 2)

[1494] 2. The system of claim 1, further comprising means for the artificial intelligence engine to analyze video frames and audio data and detect inappropriate language using a specific algorithm.

[1495] (Claim 3)

[1496] 10. The system of claim 1, further comprising means for recording timestamps and scene tags of frames containing profanity and displaying them on a user's dashboard.

[1497] "Example 2: Combining Emotion Engines"

[1498] (Claim 1)

[1499] a means for a user to select or upload a video they wish to view;

[1500] a server for storing and analyzing video files;

[1501] An artificial intelligence engine that extracts the analyzed video content as a feature vector and identifies scenes containing inappropriate content;

[1502] an emotion recognition engine for collecting user emotion data and associating the data with inappropriate scenes;

[1503] A means to tag inappropriate scenes identified by the artificial intelligence engine and emotion recognition engine and notify them on the parent dashboard;

[1504] A means for parents to review and approve corrections of inappropriate scenes;

[1505] an editing module for converting inappropriate scenes into appropriate content;

[1506] A means for generating an edited video and providing it to a user terminal;

[1507] A system including:

[1508] (Claim 2)

[1509] 2. The system of claim 1, further comprising means for the artificial intelligence engine to analyze video frames and audio data and detect inappropriate language using a specific algorithm.

[1510] (Claim 3)

[1511] 10. The system of claim 1, further comprising means for recording timestamps and scene tags of frames containing profanity and associating them with emotion data generated by an emotion recognition engine.

[1512] "Application example 2 when combining emotion engines"

[1513] (Claim 1)

[1514] a means for a user to select or upload digital content that they wish to view;

[1515] a server for storing and analyzing digital content files;

[1516] a machine learning engine for extracting the analyzed digital content as a feature vector and identifying scenes containing inappropriate content;

[1517] A means to tag inappropriate scenes identified by the machine learning engine and notify them in a parental interface;

[1518] A means for parents to review and approve corrections of inappropriate scenes;

[1519] an editing module for converting inappropriate scenes into appropriate content;

[1520] a means for collecting user emotion data and associating the data with inappropriate scenes using an emotion engine;

[1521] A means of utilizing emotional data to identify inappropriate scenes using an artificial intelligence engine; and

[1522] A means of collecting and learning from emotional data when parents approve or reject modifications;

[1523] means for generating edited digital content and providing it to a user terminal;

[1524] A system including:

[1525] (Claim 2)

[1526] 10. The system of claim 1, further comprising a machine learning engine for analyzing frames and audio data of digital content and detecting profanity using a specific algorithm.

[1527] (Claim 3)

[1528] 10. The system of claim 1, further comprising means for recording a timestamp and a scene tag of a frame containing profanity. [Explanation of symbols]

[1529] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for a user to select or upload a video they wish to view; a server for storing and analyzing video files; An artificial intelligence engine that extracts the analyzed video content as a feature vector and identifies scenes containing inappropriate content; A means to tag inappropriate scenes identified by the AI ​​engine and notify them on the parent dashboard; A means for parents to review and approve corrections of inappropriate scenes; an editing module for converting inappropriate scenes into appropriate content; A means for generating an edited video and providing it to a user terminal; A system including:

2. 10. The system of claim 1, further comprising means for the artificial intelligence engine to analyze video frames and audio data and detect profanity using a specific algorithm.

3. 2. The system of claim 1, further comprising means for recording a timestamp and a scene tag of a frame containing profanity.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A

Cited By

  • Content delivery system and device

    JP7858876B1