System
A system using voice and image recognition technologies automates video compliance checks, addressing the inefficiencies of manual review by issuing tokens for compliant videos and providing feedback for non-compliant content, thus enhancing regulatory compliance and reducing costs.
Patent Information
- Application Number
- JP2024116550
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
The expansion of the video distribution market has led to the challenge of manually reviewing cut-out videos for regulatory compliance, which is time-consuming and costly, and creators often struggle to ensure their content does not contain inappropriate or copyright-infringing material.
A system that utilizes voice recognition technology to convert audio into text data and image recognition technology to analyze video content, determining compliance with regulations, and issues authentication tokens for compliant videos while providing notifications for non-compliant content.
This system allows creators to efficiently check video compliance with regulations, reducing management costs and ensuring platforms can distribute content quickly and securely.
Smart Images

Figure 2026015076000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] With the expansion of the video distribution market, many creators are creating cut-out videos, but manual review to ensure these videos comply with regulations is time-consuming and costly. Furthermore, if some cut-out videos contain inappropriate content or copyright infringement, this could have a negative impact on distributors and platforms. Furthermore, it is often difficult for creators to complete the review process themselves, so a system for managing these processes appropriately and efficiently is needed. [Means for solving the problem]
[0005] The present invention provides a means for receiving a video file and converting the audio portion into text data using voice recognition technology, a means for analyzing the generated text data by comparing it with pre-defined regulations, and a means for analyzing the video portion of the video using image recognition technology. This allows for the creation of a system that includes a means for determining whether a video is permitted based on the results of the analysis of the text data and video data, and for issuing a token if the video is permitted. The token is sent to the user's device and can be pasted in the description section of the permitted video. The system also includes a means for sending a notification to the user's device if the analysis results in "review" or "denial." This system allows creators to easily and efficiently check whether their videos comply with regulations, allowing distributors and platforms to provide sound content while reducing management costs.
[0006] A "video file" is a digital data file that contains video and audio and is in a format that can be distributed and played.
[0007] "Analysis" refers to the process of analyzing input data and understanding its content and characteristics.
[0008] "Voice recognition technology" refers to the technology that analyzes voice data and converts it into text data.
[0009] "Text data" is character string data generated by speech recognition technology, and is a written representation of the contents of speech.
[0010] "Regulations" refer to pre-established rules and standards regarding video and audio content.
[0011] "Image recognition technology" refers to the technology of analyzing video data and recognizing objects and text within it.
[0012] A "token" is authentication data issued to an authorized video that indicates that the video complies with certain regulations.
[0013] "User terminal" refers to a device used by a user, such as a computer, smartphone, or tablet.
[0014] "Notification" refers to information or messages sent from the system to the user. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[0037] The automated video review AI system receives video files uploaded by users and analyzes the audio and video data to determine whether they comply with pre-set regulations. The specific program processing is explained below.
[0038] First, the user uploads the edited video file from the device to the platform, and the device then sends the video file to the server, completing the video upload.
[0039] The server receives the uploaded video file and begins analyzing the contents of the video file. Specifically, it converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is obtained as a string of characters.
[0040] The server then analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure it does not contain inappropriate language or copyright-infringing content.
[0041] At the same time, the server uses image recognition technology to analyze the video portion of the video file, extracting specific frames to determine whether the video contains any content that violates regulations, such as copyrighted images or inappropriate content.
[0042] Based on the results of these analyses, the server determines whether the video is permitted. If permitted, the server issues an authentication token for the video. This token certifies that the video complies with the regulations and is stored in a database in association with the video ID.
[0043] The server sends the issued token to the user's device, and the user can paste the token into the description of the video to indicate that the video is appropriate for viewers.
[0044] On the other hand, if the analysis result is "review" or "rejection," the server will send a notification to the user's device. The user can check the notification, make any necessary corrections, and then upload the video again for review.
[0045] As a concrete example, consider the case where a user creates a video by clipping interesting moments from a live broadcast and uploads the video. The server analyzes the audio of the video to check, for example, whether it contains inappropriate language. It also analyzes the video portion to check whether it contains copyright infringement or inappropriate content. If the analysis results are satisfactory, the server issues an authorization token and notifies the user. The user can indicate the integrity of the video by pasting this token in the video's description.
[0046] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations, and supports distributors and platforms in providing sound content.
[0047] The processing flow will be explained below.
[0048] Step 1:
[0049] After the user has finished editing the video, they open the platform's upload screen, click the upload button, select the video file, and click the "Upload" button.
[0050] Step 2:
[0051] The device sends the video file specified by the user to the server. Once the sending is complete, a confirmation message indicating that the upload is complete is displayed on the screen.
[0052] Step 3:
[0053] The server receives the video file sent from the device, checks the video file structure (format and length) and verifies that there are no errors, generates a video ID and stores it in the video database.
[0054] Step 4:
[0055] The server passes the audio portion of the received video file to a speech recognition module to extract the audio data, converts the extracted audio data into text (STT: Speech-to-Text), and temporarily stores the converted text data.
[0056] Step 5:
[0057] The server passes the generated text data to a natural language processing module, which performs profanity filtering and copyright infringement checks, judges the analysis results based on regulations, and records the results.
[0058] Step 6:
[0059] The server passes the video portion of the video file to an image recognition module, which extracts specific frames from the video data and analyzes them for content that violates regulations (e.g., logos, inappropriate content), compares the analysis results with regulations, and records the results.
[0060] Step 7:
[0061] The server combines the results of the text and video analysis to make a comprehensive decision. Based on the decision, if permission is granted, a token is generated. The token is associated with the video ID and stored in a database, and if a re-examination is required, this is recorded.
[0062] Step 8:
[0063] The server sends the review result to the user's device. The result will include either "permit," "review," or "deny." If the request is approved, a URL containing a token will be provided.
[0064] Step 9:
[0065] The terminal receives the examination results sent from the server, displays a notification, and displays a screen where the user can check the examination results.
[0066] Step 10:
[0067] For videos that are allowed, the user pastes the token provided by the server into the summary section of the video, then copies the token into the video description and makes it available to viewers.
[0068] Example 1
[0069] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0070] On online platforms, quickly and accurately determining the appropriateness of video content uploaded by users requires significant resources. Furthermore, the distribution of inappropriate content can undermine the platform's credibility. Traditional methods require manual review, which is labor-intensive and time-consuming. This increases the workload for content review and makes it difficult to distribute content in real time.
[0071] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0072] In this invention, the server includes means for receiving and analyzing video data, means for converting the audio portion of the video data into text data using voice recognition technology, means for analyzing the text data by checking it against preset rules, means for analyzing the video portion of the video data using image recognition technology, means for determining whether the video is permitted based on the results of the analysis of the text data and video data, and means for issuing authentication information and transmitting the authentication information to the user terminal. This allows video content uploaded by users to be quickly and accurately screened, and only appropriate content can be published on the platform.
[0073] "Video data" means media in digital form that contains visual and audio information.
[0074] "Speech recognition technology" is a technology for analyzing voice data and converting it into corresponding text data.
[0075] "Character data" is text-format data generated by voice recognition technology.
[0076] "Regulations" are pre-defined standards or conditions for determining the appropriateness of video content.
[0077] "Image recognition technology" is a technology for analyzing digital images and identifying and recognizing their contents.
[0078] "Authentication information" is digital information issued by a server to prove the authorization of video content.
[0079] "User terminal" means an electronic device used to upload videos and receive authentication information.
[0080] The following describes in detail the embodiments of the present invention. The AI system for automatic video screening receives video data uploaded by users and analyzes the audio and video data to determine whether the data conforms to pre-set regulations. Specific program processing is described below.
[0081] First, the user uploads the video data they have edited from their device to the platform. The device then sends this video data to the server, completing the video upload. The server receives the uploaded video data and begins analyzing its content. Specifically, it converts the audio portion of the video data into text data using voice recognition technology. Through this process, all of the audio information in the video is obtained as a string of characters.
[0082] The server then analyzes the generated text data and verifies whether it complies with established standards. For example, it uses profanity filters and copyright infringement checks to ensure that it does not contain inappropriate language or copyright-infringing content. At the same time, the server uses image recognition technology to analyze the video portion of the video data. This analysis involves extracting specific frames and checking whether the video contains content that violates standards. For example, it checks whether the video contains copyrighted images or inappropriate content.
[0083] Based on these analysis results, the server determines whether the video is permitted. If permission is granted, the server issues authentication information for the video. This authentication information proves that the video complies with the regulations and is stored in a database in association with the video ID. The server then sends the issued authentication information to the user's device. The user can paste the received authentication information into the video's description to indicate to viewers that the video is appropriate. On the other hand, if the analysis result is "review required" or "not permitted," the server sends a notification to the user's device. The user can review the notification, make any necessary corrections, and then upload the video again for review.
[0084] As a concrete example, consider the case where a user creates a video by cutting out interesting moments from a live broadcast and uploads the video. The server analyzes the audio of the video to check for inappropriate language. It also analyzes the video portion to check for copyright infringement or inappropriate content. If the analysis results are satisfactory, the server issues authentication information and notifies the user. The user can then paste this authentication information into the video's description to demonstrate the integrity of the video.
[0085] An example of a prompt for a generative AI model is: "I want to design a system that allows users to upload videos, analyzes the audio and video of the videos, and automatically reviews whether they comply with regulations. The server will perform this analysis using voice recognition and image recognition technology. Please tell me the specific processing steps, the technologies used, and the detailed operations at each step."
[0086] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations, and supports distributors and platforms in providing sound content.
[0087] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0088] Step 1:
[0089] The user uploads the video data after editing from the device to the platform.
[0090] Input: Edited video data
[0091] Output: Notification of completion of upload to the platform
[0092] What it does: After a user completes editing in video editing software (e.g., Adobe Premiere Pro), they open a web browser and use the platform's upload function, which selects and uploads a video file from their local disk.
[0093] Step 2:
[0094] The device sends video data to the server.
[0095] Input: User uploaded video data
[0096] Output: Data transfer to server completed
[0097] Specific operation: When the upload button is clicked, the device sends the video file to the server via an HTTP POST request. An internet connection is required for transmission.
[0098] Step 3:
[0099] The server receives the video data and stores it in storage.
[0100] Input: Video data sent from the device
[0101] Output: Video file saved in storage
[0102] Specific operation: The server saves the received video file to the specified storage (e.g., Amazon S3 bucket). Once saving is complete, metadata is registered in the database.
[0103] Step 4:
[0104] The server converts the audio portion of the video data into text data using voice recognition technology.
[0105] Input: Video file saved in storage
[0106] Output: Text data corresponding to the audio data
[0107] What happens: The server analyzes the video file, extracts the audio, and then converts the audio data into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API).
[0108] Step 5:
[0109] The server parses the generated character data against the rules.
[0110] Input: Text data generated by speech recognition technology
[0111] Output: Judgment result indicating whether the specification is met
[0112] What happens: The server analyzes the text data and performs profanity filters and copyright infringement checks to ensure that it does not contain inappropriate or infringing content.
[0113] Step 6:
[0114] The server analyzes the video portion of the video data using image recognition technology.
[0115] Input: Video file saved in storage
[0116] Output: Analysis results based on video data
[0117] Specific operation: The server extracts specific frames from the video and analyzes the video data using image recognition technology (e.g., OpenCV or Google Cloud Vision API). It checks whether the video contains any illegal content.
[0118] Step 7:
[0119] The server determines whether the video is permitted based on the results of analyzing the text data and video data, and issues authentication information.
[0120] Input: Analysis results of text data and video data after matching
[0121] Output: Certification information if approved, disapproval or reconsideration notice
[0122] Specific operation: The server evaluates the analysis results and determines whether the video complies with the regulations. If it is approved, it generates authentication information and stores it in the database. If it is not approved or requires review, it records that fact.
[0123] Step 8:
[0124] The server sends authentication information or a re-examination or denial notice to the user's device.
[0125] Input: Authorization decision result and authentication information
[0126] Output: Notification to the user's device
[0127] Specific operation: The server sends a notification to the user's device containing authentication information (if authorized) or a notification indicating reconsideration or denial. Notification methods include email and / or in-app notifications.
[0128] Step 9:
[0129] The user reviews the notification they received, makes any necessary corrections, and requests a reconsideration.
[0130] Input: Reconsideration or Denial Notice
[0131] Output: Corrected video data
[0132] What to do: The user will review the notification, make the necessary changes using video editing software, re-upload the video, and request a reconsideration.
[0133] (Application example 1)
[0134] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0135] On conventional video distribution platforms, manually reviewing video audio and video for regulatory compliance required a significant amount of time and effort, making the process highly inefficient. It was also difficult to provide users with prompt feedback, resulting in delays in video release. Furthermore, the platform lacked sufficient guidance and token issuance functions to verify the regulatory compliance of videos uploaded by users.
[0136] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0137] In this invention, the server includes means for receiving and analyzing video files, means for converting the audio portion of the video file into text data using voice recognition technology, means for analyzing the text data by comparing it with preset regulations, means for analyzing the video portion of the video file using image recognition technology, means for determining whether the video is permitted based on the analysis results of the text data and video data and issuing a token, and means for notifying the user terminal of the analysis result and the permission token and providing guidance on re-uploading. This automates video review and enables rapid feedback, allowing users to efficiently upload videos to the video distribution platform and granting tokens to permitted videos.
[0138] A "video file" is a digital file containing video and audio, and is content uploaded by a user.
[0139] "Analysis" refers to the process of individually analyzing the audio and video data of a video file to verify compliance with regulations.
[0140] "Voice recognition technology" is a technology that converts voice data into text data, and is used to obtain voice information within a video as a string of characters.
[0141] "Text data" refers to the text information of the audio data in a video that has been converted using voice recognition technology.
[0142] "Regulations" are rules and standards set in advance by video distribution platforms to evaluate whether videos comply.
[0143] "Verification" is the process of comparing text and video data with regulations to confirm compliance.
[0144] "Image recognition technology" is a technology that analyzes video data and identifies specific objects or scenes.
[0145] "Analysis results" are data that indicates whether the audio and video data conforms to regulations after analysis.
[0146] "Approved" means that the video is deemed to comply with the regulations and is allowed to be made public.
[0147] A "token" is a digital credential that proves a video complies with regulations.
[0148] "Re-uploading" refers to the act of a user re-uploading a video that has been edited to the platform.
[0149] A "guide" is an instruction or advice that tells the user what modifications are needed.
[0150] The present invention relates to an AI system for automatically reviewing videos. This system receives video files uploaded by users and analyzes their audio and video data to determine whether they comply with pre-defined regulations. Specific methods for implementing the system of the present invention are described below.
[0151] First, a user uploads the edited video file to the platform from a device such as a smartphone. The device then sends the video file to the server, completing the video upload.
[0152] The server receives the uploaded video file and begins analyzing the content of the video file. Specifically, it uses the following hardware and software:
[0153] 1. Hardware
[0154] server
[0155] Smartphone (iOS or Android)
[0156] 2. Software
[0157] Python Program
[0158] Audio Speech Recognition (ASR)
[0159] Image recognition technology (computer vision)
[0160] Regulation Check System
[0161] Next, the server converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is captured as a string of characters. For example, the server uses ASR technology to convert speech such as "Today was fun! Please come visit again." into text.
[0162] The server analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure that it does not contain inappropriate language or copyright-infringing content.
[0163] At the same time, the server uses image recognition technology to analyze the video portion of the video file. This analysis involves extracting specific frames to determine whether the video contains any content that violates regulations. For example, it checks for copyrighted images or inappropriate content. The server also performs facial recognition on the video portion to check for the presence of famous characters or specific logos.
[0164] Based on the results of these analyses, the server determines whether the video is permitted. If permission is granted, the server issues an authorization token for the video. This token proves that the video complies with the regulations and is stored in a database in association with the video ID. The server notifies the user of the issued token to their device. Once the user receives it, they can paste the token into the video's description.
[0165] On the other hand, if the analysis result is "review required" or "rejected," the server will send a notification to the user's device. The notification will contain detailed information about the problem and the necessary corrections, so the user can make corrections based on that information. If the user re-uploads the file, it can be reviewed again.
[0166] Example prompt sentence:
[0167] "Extract audio from uploaded videos, transcribe it, and check for specific keywords or phrases."
[0168] "Analyze video frames to see if they contain specific images or logos."
[0169] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations, and supports distributors and platforms in providing sound content.
[0170] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0171] Step 1:
[0172] Users upload the edited video file from their device to the platform.
[0173] Specific operation: A user opens the smartphone app, selects a video file, and taps the upload button, which causes the device to send the video file to the server.
[0174] Input: User's video file
[0175] Output: Video file sent to the server
[0176] Step 2:
[0177] The server receives the uploaded video file and begins analyzing the audio portion of the video file.
[0178] Specific operation: The server extracts the audio data from the video file and converts it into text data using speech recognition technology.
[0179] Input: Video file (audio data)
[0180] Output: Text data of the audio portion
[0181] Step 3:
[0182] The server analyzes the generated text data and checks whether it complies with pre-set regulations.
[0183] Specific operation: The server verifies the text data using a profanity filter and copyright infringement checking system.
[0184] Input: Text data of the audio portion
[0185] Output: Analysis results of text data (match / non-match)
[0186] Step 4:
[0187] The server analyzes the video portion of the video file, extracts specific frames and uses image recognition technology.
[0188] How it works: The server extracts frames from the video and uses computer vision technology to check for facial recognition and the presence of specific logos.
[0189] Input: Video file (video data)
[0190] Output: Analysis results of video data (suitable / unsuitable)
[0191] Step 5:
[0192] The server determines whether the video complies with regulations based on the results of analyzing the audio and video data.
[0193] Specific operation: The server integrates the analysis results and makes a final decision on whether to allow or deny the request.
[0194] Input: Analysis results of audio data, analysis results of video data
[0195] Output: Video approval / disapproval
[0196] Step 6:
[0197] If the video is authorized, the server issues an authorization token for the video.
[0198] What happens: The server generates a unique authentication token for each authorized video and stores it in a database.
[0199] Input: Video permission decision
[0200] Output: Authorization token
[0201] Step 7:
[0202] The server notifies the user terminal of the issued token and provides guidance on how to re-upload.
[0203] Specific operation: The server sends a notification to the user terminal, displays the authorization token along with the analysis result, or provides a correction guide in case of reconsideration or denial.
[0204] Input: Token issuance result
[0205] Output: Notification to user device (token or correction guide)
[0206] Step 8:
[0207] The user can check the notification, make any necessary corrections, and then upload the video again.
[0208] Specific actions: The user will check the notification, make the necessary corrections, and re-upload the video to the platform.
[0209] Input: Notification from the server
[0210] Output: Corrected video file
[0211] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0212] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[0213] The AI video review system receives video files uploaded by users and analyzes audio and video data to determine whether they comply with pre-defined regulations. Furthermore, the invention incorporates an emotion engine that recognizes user emotions, assisting the analysis process and improving the user experience.
[0214] First, the user uploads the edited video file from their device to the platform, and the device sends the video file to the server, displaying a confirmation message on the screen that the upload is complete.
[0215] The server receives the uploaded video file and begins analyzing the contents of the video file. Specifically, it converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is obtained as a string of characters.
[0216] The server then analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure it does not contain inappropriate language or copyright-infringing content.
[0217] At the same time, the server uses image recognition technology to analyze the video portion of the video file, extracting specific frames to determine whether the video contains any content that violates regulations, such as copyrighted images or inappropriate content.
[0218] Furthermore, the system incorporates an emotion engine that analyzes the user's facial expressions and tone of voice to determine their emotions, allowing the system to understand how the analysis results affect the user and provide more appropriate feedback.
[0219] The emotion engine records the user's emotional data and provides feedback based on the analysis. This data is also used as auxiliary information when determining whether a video complies with regulations. For example, if the user is feeling anxious or nervous, the system will carefully convey the analysis results, taking that emotion into consideration.
[0220] Based on the analysis, the server determines whether the video is permitted. If permitted, the server issues an authentication token for the video. This token certifies that the video complies with the regulations and is stored in a database in association with the video ID.
[0221] The server sends the issued token to the user's device, and the user can paste the token into the description of the video to indicate that the video is appropriate for viewers.
[0222] On the other hand, if the analysis result is "review required" or "rejected," a notification is sent to the user's device. The emotion engine recognizes the user's emotions at this time and selects an appropriate notification message to alleviate discomfort or stress, for example. The user can review the notification and make any necessary corrections before uploading the video for review again.
[0223] As a concrete example, consider the case where a user creates a video by cutting out interesting moments from a live broadcast and uploads it. The server analyzes the audio of the video to check for inappropriate language, for example. It also analyzes the video portion to ensure that it does not contain copyright infringement or inappropriate content. The emotion engine monitors the user's emotions when receiving the video review results and provides appropriate feedback. If the analysis results are satisfactory, the server issues an authorization token and notifies the user. The user can paste this token in the video description to indicate the integrity of the video to viewers.
[0224] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations and supports distributors and platforms in providing healthy content. Furthermore, the incorporation of an emotion engine improves the user experience and makes the analysis process run more smoothly.
[0225] The processing flow will be explained below.
[0226] Step 1:
[0227] After the user has finished editing the video, they open the platform's upload screen, click the upload button, select the video file, and click the "Upload" button.
[0228] Step 2:
[0229] The device sends the video file specified by the user to the server. Once the sending is complete, a confirmation message indicating that the upload is complete is displayed on the screen.
[0230] Step 3:
[0231] The server receives the received video file, checks the structure (format and length) of the video file, and verifies that there are no errors. If there are no problems, it generates a video ID and stores it in the video database.
[0232] Step 4:
[0233] The server extracts the audio portion of the video file and converts the audio data into text data using speech recognition technology (STT technology). The converted text data is temporarily stored.
[0234] Step 5:
[0235] The server passes the generated text data to a natural language processing module, which analyzes the text data by checking it against pre-defined regulations, performs profanity filters and copyright infringement checks, and records the analysis results.
[0236] Step 6:
[0237] The server extracts the video portion of the video file and analyzes it using image recognition technology. It extracts specific frames and checks whether they contain any content that violates regulations (e.g., logos, inappropriate content). It then compares the analysis results with regulations and records the results.
[0238] Step 7:
[0239] The server integrates the analysis results and makes a comprehensive decision. Based on this, it determines whether the video is permitted, and if so, generates a token. The token is associated with the video ID and stored in the database.
[0240] Step 8:
[0241] The server uses an emotion engine to analyze the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to determine the user's emotions. The analysis results are recorded and reflected in feedback.
[0242] Step 9:
[0243] The server then takes into account the user's sentiment analysis results and sends a token to the user's device, which contains information indicating that the video complies with the regulations.
[0244] Step 10:
[0245] The terminal receives the audit result and token sent from the server, displays a notification, and displays a screen where the user can check the audit result and token.
[0246] Step 11:
[0247] For videos that are allowed, the user pastes the token provided by the server into the summary section of the video, then copies the token into the video description and makes it available to viewers.
[0248] Step 12:
[0249] If the analysis result is "review" or "rejection," the server sends a notification to the user's device. At this time, the emotion engine monitors the user's emotions and provides appropriate feedback. If necessary, the emotion engine also makes suggestions on how to correct the video. The user checks the notification, makes any necessary corrections, and then uploads the video again for review.
[0250] Example 2
[0251] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0252] With the increase in video content, it is becoming increasingly important to screen videos for inappropriate content and content that infringes copyright. However, manual screening is time-consuming and labor-intensive, and it is often difficult to take appropriate action. In addition, it is necessary to consider the user's feelings regarding the screening results, so a method for performing these tasks automatically is needed.
[0253] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0254] In this invention, the server includes means for receiving and analyzing video files, means for converting the audio portion of the video file into text data using voice recognition technology, means for analyzing the text data by comparing it with pre-set regulations, means for analyzing the video portion of the video file using image recognition technology, means for assisting the analysis process with an emotion engine that recognizes user emotions, means for determining whether the video is permitted based on the analysis results of the text data and video data, and means for issuing an authentication token, and means for transmitting the authentication token to the user terminal. This makes it possible to efficiently check whether the video complies with regulations and improve the user experience.
[0255] Key Word Definitions
[0256] A "video file" is a file format for digital data that contains video and audio.
[0257] "Receiving" refers to the server taking in digital data sent from a user terminal.
[0258] "Analysis" refers to the extraction and evaluation of detailed information from the audio and video data of a video file using specific algorithms and techniques.
[0259] "Speech recognition technology" is a technology for converting voice data into text data.
[0260] "Text data" is data in the form of a string of characters converted using speech recognition technology.
[0261] "Regulations" are pre-established rules and standards for determining the suitability of video content.
[0262] "Verification" is the process of comparing the acquired text and video data with the regulations to see if they match.
[0263] "Image recognition technology" is a technology that uses specific algorithms to identify and analyze objects and scenes in video data.
[0264] An "emotion engine" is a technology that reads emotions from a user's facial expressions, tone of voice, etc., and assists in the analysis process.
[0265] An "authentication token" is electronic proof data issued to prove that a video complies with regulations.
[0266] A "terminal" is an electronic device such as a computer, smartphone, or tablet that is operated by a user.
[0267] "Notification" refers to the act of sending analysis results and other information from the server to the user terminal.
[0268] A "server" is a central processing unit that receives video files, analyzes them, issues authentication tokens, and so on.
[0269] "Feedback" refers to the method and content of the analysis results and notifications provided by the server to the user.
[0270] MODE FOR CARRYING OUT THE INVENTION
[0271] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[0272] System Overview
[0273] The AI video automatic review system receives video files uploaded by users and analyzes audio and video data to determine whether they comply with pre-defined regulations. Furthermore, the present invention also incorporates an emotion engine that recognizes user emotions, assisting the analysis process and improving the user experience.
[0274] Uploading and receiving video files
[0275] First, the user uploads the edited video file from their device to the platform. The device then sends the video file to the server, and a confirmation message appears on the screen confirming the upload. The server then receives the uploaded video file and begins to analyze its contents.
[0276] Audio analysis
[0277] Specifically, the server converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is captured as a string of characters. This process uses speech recognition software such as the Google Cloud Speech-to-Text API.
[0278] Text data regulation check
[0279] The server then analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure it does not contain inappropriate language or copyright-infringing content.
[0280] Video Analysis
[0281] At the same time, the server analyzes the video portion of the video file using image recognition technology. This involves extracting specific frames and identifying whether the video contains any content that violates regulations, such as copyrighted images or inappropriate content. This process is performed using image recognition software such as Amazon Rekognition.
[0282] Emotion engine assists the analysis process
[0283] Furthermore, the system incorporates an emotion engine that analyzes the user's facial expressions and tone of voice to determine their emotions. This allows the system to understand how the analysis results affect the user and provide more appropriate feedback. The emotion engine uses the emotion recognition API of Azure Cognitive Services.
[0284] Notification of analysis results and response
[0285] Based on the analysis results, the server determines whether the video is permitted. If permission is granted, the server issues an authentication token for the video. This token certifies that the video complies with the regulations and is stored in a database in association with the video ID. The server then sends the issued token to the user's device.
[0286] On the other hand, if the analysis result is "review required" or "rejected," a notification is sent to the user's device. The emotion engine recognizes the user's emotions at this time and selects an appropriate notification message to alleviate discomfort or stress. The user can review the notification, make any necessary corrections, and then upload the video again for review.
[0287] Examples of specific examples and prompts
[0288] As a concrete example, consider the case where a user creates a video by cutting out interesting moments from a live broadcast and uploads it. The server analyzes the audio of the video to check for inappropriate language, for example. It also analyzes the video portion to ensure that it does not contain copyright infringement or inappropriate content. The emotion engine monitors the user's emotions when receiving the video review results and provides appropriate feedback. If the analysis results are satisfactory, the server issues an authorization token and notifies the user. The user can paste this token in the video description to indicate the integrity of the video to viewers.
[0289] Prompt Sentence Examples
[0290] "I would like to upload a video file for review. Please make sure that the analysis result does not contain any inappropriate language or copyright infringement. Also, please monitor my emotions when receiving the review result and provide appropriate feedback."
[0291] This system allows users to efficiently check whether videos comply with regulations, enabling broadcasters and platforms to provide healthy content. The built-in emotion engine also improves the user experience and makes the analysis process smoother.
[0292] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0293] Program processing flow
[0294] Step 1: Upload your video file
[0295] The user uploads the edited video file from the device to the platform. As input, there is a video file previously edited by the user, which the device sends to the server. The device transfers the video file to the server and displays the upload progress with a progress bar. As output, the server receives the video file and sends a confirmation message to the device that the upload is complete.
[0296] Step 2: Receiving the video file and preparing it for analysis
[0297] The server receives the uploaded video file and begins preparations for analysis. The input is the video file stored on the server side. Specifically, the server determines where to save the received video file and performs preparations for analysis (storing the file, initializing the necessary analysis modules, etc.). The output is a video file that is ready for analysis.
[0298] Step 3: Audio analysis
[0299] The server extracts the audio portion of the video file and converts the audio data to text using speech recognition technology. The input is the audio data extracted from the video file. For example, the server extracts the audio track using a tool such as ffmpeg and sends the audio data to the Google Cloud Speech-to-Text API. The output is the audio data converted to text.
[0300] Step 4: Regulation check of text data
[0301] The server analyzes the generated text data and checks it against pre-set regulations. The input is text data obtained through speech analysis. Specifically, the server applies a profanity filter to the text data using regular expressions to check for inappropriate words. It then performs text mining to check for the inclusion of words related to copyright. The output is the result of the regulation check.
[0302] Step 5: Video Analysis
[0303] The server extracts the video portion of the video and analyzes the video using image recognition technology. The input is the video data from the video file. For example, the server uses tools such as Amazon Rekognition to extract specific frames and check for violations. The output is the analysis results of the video data.
[0304] Step 6: Sentiment Analysis
[0305] The server uses an emotion engine to analyze the user's emotions. As input, data related to the user's emotions (e.g., facial expressions and tone of voice) is provided. Specifically, the server sends the user's profile picture and video thumbnails to the emotion recognition API to obtain emotion data. As output, the user's emotion data is obtained.
[0306] Step 7: Judging and notifying analysis results
[0307] The server determines whether to allow a video based on the analysis results of the text data and video data. The input is the analysis results of the text data and video data. If permission is granted, the server generates an authentication token, stores it in a database, and sends it to the user's device. The output is an authentication token and a notification. If permission is denied or re-examination is required, the server generates an appropriate notification message and sends a message using an emotion engine that minimizes the impact on the user. The input is the analysis results and the user's emotion data. As a specific example of operation, the server determines the content of the notification message and sends it to the user's device. The output is a notification of denial or re-examination.
[0308] This ensures that the entire process, from uploading the video file to analysis and notification of the results, proceeds smoothly, and appropriate feedback is provided to the user.
[0309] (Application example 2)
[0310] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0311] In recent years, the distribution of video content has rapidly increased, making the integrity of the content and compliance with legal regulations important issues. However, manual video review is labor-intensive and time-consuming, and has limited scalability. There is also a need for a method to reduce the anxiety and stress users feel when receiving review results. Furthermore, determining whether content is inappropriate or copyright infringing is complex, creating a demand for an accurate and fast automated review system.
[0312] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0313] In this invention, the server includes means for receiving and analyzing video files, means for converting the audio portion of the video file into text data using voice recognition technology, means for analyzing the text data by comparing it with pre-set regulations, means for analyzing the video portion of the video file using image recognition technology, means for determining whether the video is permitted based on the results of the analysis of the text data and video data and issuing a token, means for transmitting the token to a user terminal, means for the user terminal to display the analysis results and provide feedback, and means for recognizing user emotions and assisting the analysis process. This enables automatic review of the soundness of video content and compliance with laws and regulations, and also realizes the provision of feedback that takes user emotions into consideration.
[0314] A "video file" is a file containing video and audio stored in digital format.
[0315] An "analyzing means" is a device or software module for analyzing the contents of a video file and extracting specific information.
[0316] "Speech recognition technology" is a technology for converting voice data into text data.
[0317] "Text data" refers to digital data converted into character information.
[0318] "Regulations" refer to pre-established rules and standards.
[0319] "Image recognition technology" is a technology that analyzes video data and extracts specific information.
[0320] A "token" refers to authentication information issued by the system, which is a digital code that certifies a specific operation or right.
[0321] A "user terminal" is a device such as a computer or smartphone that a user uses to upload video files and receive analysis results.
[0322] A "means for providing feedback" is a device or software module for communicating analysis results or other information to a user.
[0323] The "means for recognizing emotions" is a device or software module for analyzing the user's facial expressions and voice to determine their emotional state.
[0324] The "means for assisting the analysis process" is a device or software module for adjusting the presentation method of the analysis results based on the emotion recognition results.
[0325] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[0326] The automated video review system receives video files uploaded by users and analyzes audio and video data to determine whether they comply with pre-defined regulations. Furthermore, the present invention also incorporates an emotion engine that recognizes user emotions, which assists the analysis process and improves the user experience.
[0327] The server first receives the video file uploaded from the user's device. Next, it converts the audio portion of the video file into text data using speech recognition technology. This speech recognition typically uses existing systems such as Google Speech-to-Text API or IBM Watson Speech to Text. This text data is then analyzed based on set regulations to check, for example, whether it contains inappropriate language or copyrighted content.
[0328] At the same time, the server analyzes the video portion of the video file using image recognition technology, which uses machine learning models such as OpenCV and TensorFlow, to extract specific frames of the video and check for inappropriate content or potential copyright infringement.
[0329] The emotion engine analyzes the user's facial expressions and tone of voice during the video screening to recognize their emotions. Existing services such as Face++ and Microsoft Emotion API are used for emotion recognition. The emotion engine records the user's emotional state when receiving the video screening results and adjusts the way the analysis results are presented.
[0330] If the analysis results show that the video complies with the regulations, the server issues an authentication token to the video. This token certifies that the video is legitimate and can be displayed to viewers by users pasting it into the video description. The user's device receives this token and provides a method for appropriately placing it in the video description.
[0331] On the other hand, if the analysis results in "review" or "denial," a notification is sent to the user's device. The emotion engine recognizes the user's emotions at this time and selects an appropriate notification message to alleviate any discomfort or stress. After receiving the notification, the user can make any necessary corrections and upload the video again for review.
[0332] As a concrete example, consider the case where a user creates a video recording a funny moment during a live broadcast and uploads it to the system. The server analyzes the audio of the video to check for inappropriate language. It also analyzes the video portion to check for any content that may infringe copyright. The emotion engine monitors the user's emotions when receiving the review results and provides appropriate feedback. If there are no problems with the analysis results, the server issues an authentication token for the video and notifies the user. The user can paste this token in the video description to indicate the integrity of the video to viewers.
[0333] Additionally, examples of prompt sentences for the generative AI model to explain the analysis process of this system are shown below.
[0334] Example prompt sentence:
[0335] It analyzes the input mp4 video file and checks the following based on the audio and video data:
[0336] 1. Does it contain inappropriate language (profanity filter)?
[0337] 2. Does it contain copyright infringement (detecting specific images, music, etc.)?
[0338] Once the analysis is complete, determine relevance and generate feedback. Additionally, monitor the user's emotions and provide appropriate feedback.
[0339] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0340] Step 1:
[0341] A user uploads a video file.
[0342] The user selects a video file and uploads it to the system from the device, along with the video file's metadata (e.g., file name, size, format).
[0343] Input: A video file selected by the user
[0344] Output: Uploaded video files and metadata are sent to the server.
[0345] Step 2:
[0346] The server receives the video file and prepares it for analysis.
[0347] The server temporarily stores the received video file and prepares to start the analysis process, specifically separating the audio and video data.
[0348] Input: Uploaded video file
[0349] Output: Audio and video data are separated.
[0350] Step 3:
[0351] The server converts the voice data into text data using voice recognition technology.
[0352] The server inputs the separated voice data into a voice recognition engine (for example, Google Speech-to-Text API) and generates text data from the voice.
[0353] Input: Audio data
[0354] Output: Text data
[0355] Step 4:
[0356] The server analyzes the text data and checks it against regulations.
[0357] The generated text data is compared with the set regulations (filtering inappropriate language, checking copyright infringement, etc.).
[0358] Input: Text data
[0359] Output: Analysis results based on regulations (compliance / non-compliance)
[0360] Step 5:
[0361] The server analyzes the video data using image recognition technology.
[0362] The server feeds the video data into an image recognition engine (e.g., OpenCV or TensorFlow), extracts specific frames, and checks them for inappropriate content or potential copyright infringement.
[0363] Input: Video data
[0364] Output: Analysis results based on image recognition (match / failure)
[0365] Step 6:
[0366] The emotion engine recognizes the user's emotions.
[0367] The server uses an emotion engine to analyze the user's facial expressions and tone of voice and recognize their emotional state before notifying the user of the analysis results.
[0368] Input: User's facial expression data, voice tone data
[0369] Output: User's emotional state data (e.g., relief, tension, anxiety)
[0370] Step 7:
[0371] The server generates feedback based on the analysis results and emotional data.
[0372] The server generates appropriate feedback based on the emotion recognition results and provides the analysis results to the user.
[0373] Input: Analysis results, emotional state data
[0374] Output: Feedback message provided to the user
[0375] Step 8:
[0376] The server determines the video's authorization and issues an authorization token.
[0377] If the video complies with the regulations, the server generates an authentication token and sends it to the user's device.
[0378] Input: All analysis results
[0379] Output: Authentication token
[0380] Step 9:
[0381] The server sends the token to the user terminal.
[0382] Once the authentication token is generated, the server sends it to the user's device, and the user pastes the token into the video description.
[0383] Input: Authentication Token
[0384] Output: Token sent to user device
[0385] Step 10:
[0386] The server sends a notification to the user.
[0387] If the analysis result is "re-examination" or "denial," a notification is sent to the user's device to inform them of the necessary corrections.
[0388] Input: Analysis results (review / rejection)
[0389] Output: Message to be sent to the user
[0390] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0391] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0392] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0393] [Second embodiment]
[0394] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0395] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0396] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0397] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0398] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0399] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0400] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0401] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0402] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0403] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0404] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0405] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0406] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[0407] The automated video review AI system receives video files uploaded by users and analyzes the audio and video data to determine whether they comply with pre-set regulations. The specific program processing is explained below.
[0408] First, the user uploads the edited video file from the device to the platform, and the device then sends the video file to the server, completing the video upload.
[0409] The server receives the uploaded video file and begins analyzing the contents of the video file. Specifically, it converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is obtained as a string of characters.
[0410] The server then analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure it does not contain inappropriate language or copyright-infringing content.
[0411] At the same time, the server uses image recognition technology to analyze the video portion of the video file, extracting specific frames to determine whether the video contains any content that violates regulations, such as copyrighted images or inappropriate content.
[0412] Based on the results of these analyses, the server determines whether the video is permitted. If permitted, the server issues an authentication token for the video. This token certifies that the video complies with the regulations and is stored in a database in association with the video ID.
[0413] The server sends the issued token to the user's device, and the user can paste the token into the description of the video to indicate that the video is appropriate for viewers.
[0414] On the other hand, if the analysis result is "review" or "rejection," the server will send a notification to the user's device. The user can check the notification, make any necessary corrections, and then upload the video again for review.
[0415] As a concrete example, consider the case where a user creates a video by clipping interesting moments from a live broadcast and uploads the video. The server analyzes the audio of the video to check, for example, whether it contains inappropriate language. It also analyzes the video portion to check whether it contains copyright infringement or inappropriate content. If the analysis results are satisfactory, the server issues an authorization token and notifies the user. The user can indicate the integrity of the video by pasting this token in the video's description.
[0416] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations, and supports distributors and platforms in providing sound content.
[0417] The processing flow will be explained below.
[0418] Step 1:
[0419] After the user has finished editing the video, they open the platform's upload screen, click the upload button, select the video file, and click the "Upload" button.
[0420] Step 2:
[0421] The device sends the video file specified by the user to the server. Once the sending is complete, a confirmation message indicating that the upload is complete is displayed on the screen.
[0422] Step 3:
[0423] The server receives the video file sent from the device, checks the video file structure (format and length) and verifies that there are no errors, generates a video ID and stores it in the video database.
[0424] Step 4:
[0425] The server passes the audio portion of the received video file to a speech recognition module to extract the audio data, converts the extracted audio data into text (STT: Speech-to-Text), and temporarily stores the converted text data.
[0426] Step 5:
[0427] The server passes the generated text data to a natural language processing module, which performs profanity filtering and copyright infringement checks, judges the analysis results based on regulations, and records the results.
[0428] Step 6:
[0429] The server passes the video portion of the video file to an image recognition module, which extracts specific frames from the video data and analyzes them for content that violates regulations (e.g., logos, inappropriate content), compares the analysis results with regulations, and records the results.
[0430] Step 7:
[0431] The server combines the results of the text and video analysis to make a comprehensive decision. Based on the decision, if permission is granted, a token is generated. The token is associated with the video ID and stored in a database, and if a re-examination is required, this is recorded.
[0432] Step 8:
[0433] The server sends the review result to the user's device. The result will include either "permit," "review," or "deny." If the request is approved, a URL containing a token will be provided.
[0434] Step 9:
[0435] The terminal receives the examination results sent from the server, displays a notification, and displays a screen where the user can check the examination results.
[0436] Step 10:
[0437] For videos that are allowed, the user pastes the token provided by the server into the summary section of the video, then copies the token into the video description and makes it available to viewers.
[0438] Example 1
[0439] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0440] On online platforms, quickly and accurately determining the appropriateness of video content uploaded by users requires significant resources. Furthermore, the distribution of inappropriate content can undermine the platform's credibility. Traditional methods require manual review, which is labor-intensive and time-consuming. This increases the workload for content review and makes it difficult to distribute content in real time.
[0441] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0442] In this invention, the server includes means for receiving and analyzing video data, means for converting the audio portion of the video data into text data using voice recognition technology, means for analyzing the text data by checking it against preset rules, means for analyzing the video portion of the video data using image recognition technology, means for determining whether the video is permitted based on the results of the analysis of the text data and video data, and means for issuing authentication information and transmitting the authentication information to the user terminal. This allows video content uploaded by users to be quickly and accurately screened, and only appropriate content can be published on the platform.
[0443] "Video data" means media in digital form that contains visual and audio information.
[0444] "Speech recognition technology" is a technology for analyzing voice data and converting it into corresponding text data.
[0445] "Character data" is text-format data generated by voice recognition technology.
[0446] "Regulations" are pre-defined standards or conditions for determining the appropriateness of video content.
[0447] "Image recognition technology" is a technology for analyzing digital images and identifying and recognizing their contents.
[0448] "Authentication information" is digital information issued by a server to prove the authorization of video content.
[0449] "User terminal" means an electronic device used to upload videos and receive authentication information.
[0450] The following describes in detail the embodiments of the present invention. The AI system for automatic video screening receives video data uploaded by users and analyzes the audio and video data to determine whether the data conforms to pre-set regulations. Specific program processing is described below.
[0451] First, the user uploads the video data they have edited from their device to the platform. The device then sends this video data to the server, completing the video upload. The server receives the uploaded video data and begins analyzing its content. Specifically, it converts the audio portion of the video data into text data using voice recognition technology. Through this process, all of the audio information in the video is obtained as a string of characters.
[0452] The server then analyzes the generated text data and verifies whether it complies with established standards. For example, it uses profanity filters and copyright infringement checks to ensure that it does not contain inappropriate language or copyright-infringing content. At the same time, the server uses image recognition technology to analyze the video portion of the video data. This analysis involves extracting specific frames and checking whether the video contains content that violates standards. For example, it checks whether the video contains copyrighted images or inappropriate content.
[0453] Based on these analysis results, the server determines whether the video is permitted. If permission is granted, the server issues authentication information for the video. This authentication information proves that the video complies with the regulations and is stored in a database in association with the video ID. The server then sends the issued authentication information to the user's device. The user can paste the received authentication information into the video's description to indicate to viewers that the video is appropriate. On the other hand, if the analysis result is "review required" or "not permitted," the server sends a notification to the user's device. The user can review the notification, make any necessary corrections, and then upload the video again for review.
[0454] As a concrete example, consider the case where a user creates a video by cutting out interesting moments from a live broadcast and uploads the video. The server analyzes the audio of the video to check for inappropriate language. It also analyzes the video portion to check for copyright infringement or inappropriate content. If the analysis results are satisfactory, the server issues authentication information and notifies the user. The user can then paste this authentication information into the video's description to demonstrate the integrity of the video.
[0455] An example of a prompt for a generative AI model is: "I want to design a system that allows users to upload videos, analyzes the audio and video of the videos, and automatically reviews whether they comply with regulations. The server will perform this analysis using voice recognition and image recognition technology. Please tell me the specific processing steps, the technologies used, and the detailed operations at each step."
[0456] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations, and supports distributors and platforms in providing sound content.
[0457] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0458] Step 1:
[0459] The user uploads the video data after editing from the device to the platform.
[0460] Input: Edited video data
[0461] Output: Notification of completion of upload to the platform
[0462] What it does: After a user completes editing in video editing software (e.g., Adobe Premiere Pro), they open a web browser and use the platform's upload function, which selects and uploads a video file from their local disk.
[0463] Step 2:
[0464] The device sends video data to the server.
[0465] Input: User uploaded video data
[0466] Output: Data transfer to server completed
[0467] Specific operation: When the upload button is clicked, the device sends the video file to the server via an HTTP POST request. An internet connection is required for transmission.
[0468] Step 3:
[0469] The server receives the video data and stores it in storage.
[0470] Input: Video data sent from the device
[0471] Output: Video file saved in storage
[0472] Specific operation: The server saves the received video file to the specified storage (e.g., Amazon S3 bucket). Once saving is complete, metadata is registered in the database.
[0473] Step 4:
[0474] The server converts the audio portion of the video data into text data using voice recognition technology.
[0475] Input: Video file saved in storage
[0476] Output: Text data corresponding to the audio data
[0477] What happens: The server analyzes the video file, extracts the audio, and then converts the audio data into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API).
[0478] Step 5:
[0479] The server parses the generated character data against the rules.
[0480] Input: Text data generated by speech recognition technology
[0481] Output: Judgment result indicating whether the specification is met
[0482] What happens: The server analyzes the text data and performs profanity filters and copyright infringement checks to ensure that it does not contain inappropriate or infringing content.
[0483] Step 6:
[0484] The server analyzes the video portion of the video data using image recognition technology.
[0485] Input: Video file saved in storage
[0486] Output: Analysis results based on video data
[0487] Specific operation: The server extracts specific frames from the video and analyzes the video data using image recognition technology (e.g., OpenCV or Google Cloud Vision API). It checks whether the video contains any illegal content.
[0488] Step 7:
[0489] The server determines whether the video is permitted based on the results of analyzing the text data and video data, and issues authentication information.
[0490] Input: Analysis results of text data and video data after matching
[0491] Output: Certification information if approved, disapproval or reconsideration notice
[0492] Specific operation: The server evaluates the analysis results and determines whether the video complies with the regulations. If it is approved, it generates authentication information and stores it in the database. If it is not approved or requires review, it records that fact.
[0493] Step 8:
[0494] The server sends authentication information or a re-examination or denial notice to the user's device.
[0495] Input: Authorization decision result and authentication information
[0496] Output: Notification to the user's device
[0497] Specific operation: The server sends a notification to the user's device containing authentication information (if authorized) or a notification indicating reconsideration or denial. Notification methods include email and / or in-app notifications.
[0498] Step 9:
[0499] The user reviews the notification they received, makes any necessary corrections, and requests a reconsideration.
[0500] Input: Reconsideration or Denial Notice
[0501] Output: Corrected video data
[0502] What to do: The user will review the notification, make the necessary changes using video editing software, re-upload the video, and request a reconsideration.
[0503] (Application example 1)
[0504] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0505] On conventional video distribution platforms, manually reviewing video audio and video for regulatory compliance required a significant amount of time and effort, making the process highly inefficient. It was also difficult to provide users with prompt feedback, resulting in delays in video release. Furthermore, the platform lacked sufficient guidance and token issuance functions to verify the regulatory compliance of videos uploaded by users.
[0506] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0507] In this invention, the server includes means for receiving and analyzing video files, means for converting the audio portion of the video file into text data using voice recognition technology, means for analyzing the text data by comparing it with preset regulations, means for analyzing the video portion of the video file using image recognition technology, means for determining whether the video is permitted based on the analysis results of the text data and video data and issuing a token, and means for notifying the user terminal of the analysis result and the permission token and providing guidance on re-uploading. This automates video review and enables rapid feedback, allowing users to efficiently upload videos to the video distribution platform and granting tokens to permitted videos.
[0508] A "video file" is a digital file containing video and audio, and is content uploaded by a user.
[0509] "Analysis" refers to the process of individually analyzing the audio and video data of a video file to verify compliance with regulations.
[0510] "Voice recognition technology" is a technology that converts voice data into text data, and is used to obtain voice information within a video as a string of characters.
[0511] "Text data" refers to the text information of the audio data in a video that has been converted using voice recognition technology.
[0512] "Regulations" are rules and standards set in advance by video distribution platforms to evaluate whether videos comply.
[0513] "Verification" is the process of comparing text and video data with regulations to confirm compliance.
[0514] "Image recognition technology" is a technology that analyzes video data and identifies specific objects or scenes.
[0515] "Analysis results" are data that indicates whether the audio and video data conforms to regulations after analysis.
[0516] "Approved" means that the video is deemed to comply with the regulations and is allowed to be made public.
[0517] A "token" is a digital credential that proves a video complies with regulations.
[0518] "Re-uploading" refers to the act of a user re-uploading a video that has been edited to the platform.
[0519] A "guide" is an instruction or advice that tells the user what modifications are needed.
[0520] The present invention relates to an AI system for automatically reviewing videos. This system receives video files uploaded by users and analyzes their audio and video data to determine whether they comply with pre-defined regulations. Specific methods for implementing the system of the present invention are described below.
[0521] First, a user uploads the edited video file to the platform from a device such as a smartphone. The device then sends the video file to the server, completing the video upload.
[0522] The server receives the uploaded video file and begins analyzing the content of the video file. Specifically, it uses the following hardware and software:
[0523] 1. Hardware
[0524] server
[0525] Smartphone (iOS or Android)
[0526] 2. Software
[0527] Python Program
[0528] Audio Speech Recognition (ASR)
[0529] Image recognition technology (computer vision)
[0530] Regulation Check System
[0531] Next, the server converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is captured as a string of characters. For example, the server uses ASR technology to convert speech such as "Today was fun! Please come visit again." into text.
[0532] The server analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure that it does not contain inappropriate language or copyright-infringing content.
[0533] At the same time, the server uses image recognition technology to analyze the video portion of the video file. This analysis involves extracting specific frames to determine whether the video contains any content that violates regulations. For example, it checks for copyrighted images or inappropriate content. The server also performs facial recognition on the video portion to check for the presence of famous characters or specific logos.
[0534] Based on the results of these analyses, the server determines whether the video is permitted. If permission is granted, the server issues an authorization token for the video. This token proves that the video complies with the regulations and is stored in a database in association with the video ID. The server notifies the user of the issued token to their device. Once the user receives it, they can paste the token into the video's description.
[0535] On the other hand, if the analysis result is "review required" or "rejected," the server will send a notification to the user's device. The notification will contain detailed information about the problem and the necessary corrections, so the user can make corrections based on that information. If the user re-uploads the file, it can be reviewed again.
[0536] Example prompt sentence:
[0537] "Extract audio from uploaded videos, transcribe it, and check for specific keywords or phrases."
[0538] "Analyze video frames to see if they contain specific images or logos."
[0539] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations, and supports distributors and platforms in providing sound content.
[0540] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0541] Step 1:
[0542] Users upload the edited video file from their device to the platform.
[0543] Specific operation: A user opens the smartphone app, selects a video file, and taps the upload button, which causes the device to send the video file to the server.
[0544] Input: User's video file
[0545] Output: Video file sent to the server
[0546] Step 2:
[0547] The server receives the uploaded video file and begins analyzing the audio portion of the video file.
[0548] Specific operation: The server extracts the audio data from the video file and converts it into text data using speech recognition technology.
[0549] Input: Video file (audio data)
[0550] Output: Text data of the audio portion
[0551] Step 3:
[0552] The server analyzes the generated text data and checks whether it complies with pre-set regulations.
[0553] Specific operation: The server verifies the text data using a profanity filter and copyright infringement checking system.
[0554] Input: Text data of the audio portion
[0555] Output: Analysis results of text data (match / non-match)
[0556] Step 4:
[0557] The server analyzes the video portion of the video file, extracts specific frames and uses image recognition technology.
[0558] How it works: The server extracts frames from the video and uses computer vision technology to check for facial recognition and the presence of specific logos.
[0559] Input: Video file (video data)
[0560] Output: Analysis results of video data (suitable / unsuitable)
[0561] Step 5:
[0562] The server determines whether the video complies with regulations based on the results of analyzing the audio and video data.
[0563] Specific operation: The server integrates the analysis results and makes a final decision on whether to allow or deny the request.
[0564] Input: Analysis results of audio data, analysis results of video data
[0565] Output: Video approval / disapproval
[0566] Step 6:
[0567] If the video is authorized, the server issues an authorization token for the video.
[0568] What happens: The server generates a unique authentication token for each authorized video and stores it in a database.
[0569] Input: Video permission decision
[0570] Output: Authorization token
[0571] Step 7:
[0572] The server notifies the user terminal of the issued token and provides guidance on how to re-upload.
[0573] Specific operation: The server sends a notification to the user terminal, displays the authorization token along with the analysis result, or provides a correction guide in case of reconsideration or denial.
[0574] Input: Token issuance result
[0575] Output: Notification to user device (token or correction guide)
[0576] Step 8:
[0577] The user can check the notification, make any necessary corrections, and then upload the video again.
[0578] Specific actions: The user will check the notification, make the necessary corrections, and re-upload the video to the platform.
[0579] Input: Notification from the server
[0580] Output: Corrected video file
[0581] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0582] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[0583] The AI video review system receives video files uploaded by users and analyzes audio and video data to determine whether they comply with pre-defined regulations. Furthermore, the invention incorporates an emotion engine that recognizes user emotions, assisting the analysis process and improving the user experience.
[0584] First, the user uploads the edited video file from their device to the platform, and the device sends the video file to the server, displaying a confirmation message on the screen that the upload is complete.
[0585] The server receives the uploaded video file and begins analyzing the contents of the video file. Specifically, it converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is obtained as a string of characters.
[0586] The server then analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure it does not contain inappropriate language or copyright-infringing content.
[0587] At the same time, the server uses image recognition technology to analyze the video portion of the video file, extracting specific frames to determine whether the video contains any content that violates regulations, such as copyrighted images or inappropriate content.
[0588] Furthermore, the system incorporates an emotion engine that analyzes the user's facial expressions and tone of voice to determine their emotions, allowing the system to understand how the analysis results affect the user and provide more appropriate feedback.
[0589] The emotion engine records the user's emotional data and provides feedback based on the analysis. This data is also used as auxiliary information when determining whether a video complies with regulations. For example, if the user is feeling anxious or nervous, the system will carefully convey the analysis results, taking that emotion into consideration.
[0590] Based on the analysis, the server determines whether the video is permitted. If permitted, the server issues an authentication token for the video. This token certifies that the video complies with the regulations and is stored in a database in association with the video ID.
[0591] The server sends the issued token to the user's device, and the user can paste the token into the description of the video to indicate that the video is appropriate for viewers.
[0592] On the other hand, if the analysis result is "review required" or "rejected," a notification is sent to the user's device. The emotion engine recognizes the user's emotions at this time and selects an appropriate notification message to alleviate discomfort or stress, for example. The user can review the notification and make any necessary corrections before uploading the video for review again.
[0593] As a concrete example, consider the case where a user creates a video by cutting out interesting moments from a live broadcast and uploads it. The server analyzes the audio of the video to check for inappropriate language, for example. It also analyzes the video portion to ensure that it does not contain copyright infringement or inappropriate content. The emotion engine monitors the user's emotions when receiving the video review results and provides appropriate feedback. If the analysis results are satisfactory, the server issues an authorization token and notifies the user. The user can paste this token in the video description to indicate the integrity of the video to viewers.
[0594] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations and supports distributors and platforms in providing healthy content. Furthermore, the incorporation of an emotion engine improves the user experience and makes the analysis process run more smoothly.
[0595] The processing flow will be explained below.
[0596] Step 1:
[0597] After the user has finished editing the video, they open the platform's upload screen, click the upload button, select the video file, and click the "Upload" button.
[0598] Step 2:
[0599] The device sends the video file specified by the user to the server. Once the sending is complete, a confirmation message indicating that the upload is complete is displayed on the screen.
[0600] Step 3:
[0601] The server receives the received video file, checks the structure (format and length) of the video file, and verifies that there are no errors. If there are no problems, it generates a video ID and stores it in the video database.
[0602] Step 4:
[0603] The server extracts the audio portion of the video file and converts the audio data into text data using speech recognition technology (STT technology). The converted text data is temporarily stored.
[0604] Step 5:
[0605] The server passes the generated text data to a natural language processing module, which analyzes the text data by checking it against pre-defined regulations, performs profanity filters and copyright infringement checks, and records the analysis results.
[0606] Step 6:
[0607] The server extracts the video portion of the video file and analyzes it using image recognition technology. It extracts specific frames and checks whether they contain any content that violates regulations (e.g., logos, inappropriate content). It then compares the analysis results with regulations and records the results.
[0608] Step 7:
[0609] The server integrates the analysis results and makes a comprehensive decision. Based on this, it determines whether the video is permitted, and if so, generates a token. The token is associated with the video ID and stored in the database.
[0610] Step 8:
[0611] The server uses an emotion engine to analyze the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to determine the user's emotions. The analysis results are recorded and reflected in feedback.
[0612] Step 9:
[0613] The server then takes into account the user's sentiment analysis results and sends a token to the user's device, which contains information indicating that the video complies with the regulations.
[0614] Step 10:
[0615] The terminal receives the audit result and token sent from the server, displays a notification, and displays a screen where the user can check the audit result and token.
[0616] Step 11:
[0617] For videos that are allowed, the user pastes the token provided by the server into the summary section of the video, then copies the token into the video description and makes it available to viewers.
[0618] Step 12:
[0619] If the analysis result is "review" or "rejection," the server sends a notification to the user's device. At this time, the emotion engine monitors the user's emotions and provides appropriate feedback. If necessary, the emotion engine also makes suggestions on how to correct the video. The user checks the notification, makes any necessary corrections, and then uploads the video again for review.
[0620] Example 2
[0621] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0622] With the increase in video content, it is becoming increasingly important to screen videos for inappropriate content and content that infringes copyright. However, manual screening is time-consuming and labor-intensive, and it is often difficult to take appropriate action. In addition, it is necessary to consider the user's feelings regarding the screening results, so a method for performing these tasks automatically is needed.
[0623] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0624] In this invention, the server includes means for receiving and analyzing video files, means for converting the audio portion of the video file into text data using voice recognition technology, means for analyzing the text data by comparing it with pre-set regulations, means for analyzing the video portion of the video file using image recognition technology, means for assisting the analysis process with an emotion engine that recognizes user emotions, means for determining whether the video is permitted based on the analysis results of the text data and video data, and means for issuing an authentication token, and means for transmitting the authentication token to the user terminal. This makes it possible to efficiently check whether the video complies with regulations and improve the user experience.
[0625] Key Word Definitions
[0626] A "video file" is a file format for digital data that contains video and audio.
[0627] "Receiving" refers to the server taking in digital data sent from a user terminal.
[0628] "Analysis" refers to the extraction and evaluation of detailed information from the audio and video data of a video file using specific algorithms and techniques.
[0629] "Speech recognition technology" is a technology for converting voice data into text data.
[0630] "Text data" is data in the form of a string of characters converted using speech recognition technology.
[0631] "Regulations" are pre-established rules and standards for determining the suitability of video content.
[0632] "Verification" is the process of comparing the acquired text and video data with the regulations to see if they match.
[0633] "Image recognition technology" is a technology that uses specific algorithms to identify and analyze objects and scenes in video data.
[0634] An "emotion engine" is a technology that reads emotions from a user's facial expressions, tone of voice, etc., and assists in the analysis process.
[0635] An "authentication token" is electronic proof data issued to prove that a video complies with regulations.
[0636] A "terminal" is an electronic device such as a computer, smartphone, or tablet that is operated by a user.
[0637] "Notification" refers to the act of sending analysis results and other information from the server to the user terminal.
[0638] A "server" is a central processing unit that receives video files, analyzes them, issues authentication tokens, and so on.
[0639] "Feedback" refers to the method and content of the analysis results and notifications provided by the server to the user.
[0640] MODE FOR CARRYING OUT THE INVENTION
[0641] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[0642] System Overview
[0643] The AI video automatic review system receives video files uploaded by users and analyzes audio and video data to determine whether they comply with pre-defined regulations. Furthermore, the present invention also incorporates an emotion engine that recognizes user emotions, assisting the analysis process and improving the user experience.
[0644] Uploading and receiving video files
[0645] First, the user uploads the edited video file from their device to the platform. The device then sends the video file to the server, and a confirmation message appears on the screen confirming the upload. The server then receives the uploaded video file and begins to analyze its contents.
[0646] Audio analysis
[0647] Specifically, the server converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is captured as a string of characters. This process uses speech recognition software such as the Google Cloud Speech-to-Text API.
[0648] Text data regulation check
[0649] The server then analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure it does not contain inappropriate language or copyright-infringing content.
[0650] Video Analysis
[0651] At the same time, the server analyzes the video portion of the video file using image recognition technology. This involves extracting specific frames and identifying whether the video contains any content that violates regulations, such as copyrighted images or inappropriate content. This process is performed using image recognition software such as Amazon Rekognition.
[0652] Emotion engine assists the analysis process
[0653] Furthermore, the system incorporates an emotion engine that analyzes the user's facial expressions and tone of voice to determine their emotions. This allows the system to understand how the analysis results affect the user and provide more appropriate feedback. The emotion engine uses the emotion recognition API of Azure Cognitive Services.
[0654] Notification of analysis results and response
[0655] Based on the analysis results, the server determines whether the video is permitted. If permission is granted, the server issues an authentication token for the video. This token certifies that the video complies with the regulations and is stored in a database in association with the video ID. The server then sends the issued token to the user's device.
[0656] On the other hand, if the analysis result is "review required" or "rejected," a notification is sent to the user's device. The emotion engine recognizes the user's emotions at this time and selects an appropriate notification message to alleviate discomfort or stress. The user can review the notification, make any necessary corrections, and then upload the video again for review.
[0657] Examples of specific examples and prompts
[0658] As a concrete example, consider the case where a user creates a video by cutting out interesting moments from a live broadcast and uploads it. The server analyzes the audio of the video to check for inappropriate language, for example. It also analyzes the video portion to ensure that it does not contain copyright infringement or inappropriate content. The emotion engine monitors the user's emotions when receiving the video review results and provides appropriate feedback. If the analysis results are satisfactory, the server issues an authorization token and notifies the user. The user can paste this token in the video description to indicate the integrity of the video to viewers.
[0659] Prompt Sentence Examples
[0660] "I would like to upload a video file for review. Please make sure that the analysis result does not contain any inappropriate language or copyright infringement. Also, please monitor my emotions when receiving the review result and provide appropriate feedback."
[0661] This system allows users to efficiently check whether videos comply with regulations, enabling broadcasters and platforms to provide healthy content. The built-in emotion engine also improves the user experience and makes the analysis process smoother.
[0662] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0663] Program processing flow
[0664] Step 1: Upload your video file
[0665] The user uploads the edited video file from the device to the platform. As input, there is a video file previously edited by the user, which the device sends to the server. The device transfers the video file to the server and displays the upload progress with a progress bar. As output, the server receives the video file and sends a confirmation message to the device that the upload is complete.
[0666] Step 2: Receiving the video file and preparing it for analysis
[0667] The server receives the uploaded video file and begins preparations for analysis. The input is the video file stored on the server side. Specifically, the server determines where to save the received video file and performs preparations for analysis (storing the file, initializing the necessary analysis modules, etc.). The output is a video file that is ready for analysis.
[0668] Step 3: Audio analysis
[0669] The server extracts the audio portion of the video file and converts the audio data to text using speech recognition technology. The input is the audio data extracted from the video file. For example, the server extracts the audio track using a tool such as ffmpeg and sends the audio data to the Google Cloud Speech-to-Text API. The output is the audio data converted to text.
[0670] Step 4: Regulation check of text data
[0671] The server analyzes the generated text data and checks it against pre-set regulations. The input is text data obtained through speech analysis. Specifically, the server applies a profanity filter to the text data using regular expressions to check for inappropriate words. It then performs text mining to check for the inclusion of words related to copyright. The output is the result of the regulation check.
[0672] Step 5: Video Analysis
[0673] The server extracts the video portion of the video and analyzes the video using image recognition technology. The input is the video data from the video file. For example, the server uses tools such as Amazon Rekognition to extract specific frames and check for violations. The output is the analysis results of the video data.
[0674] Step 6: Sentiment Analysis
[0675] The server uses an emotion engine to analyze the user's emotions. As input, data related to the user's emotions (e.g., facial expressions and tone of voice) is provided. Specifically, the server sends the user's profile picture and video thumbnails to the emotion recognition API to obtain emotion data. As output, the user's emotion data is obtained.
[0676] Step 7: Judging and notifying analysis results
[0677] The server determines whether to allow a video based on the analysis results of the text data and video data. The input is the analysis results of the text data and video data. If permission is granted, the server generates an authentication token, stores it in a database, and sends it to the user's device. The output is an authentication token and a notification. If permission is denied or re-examination is required, the server generates an appropriate notification message and sends a message using an emotion engine that minimizes the impact on the user. The input is the analysis results and the user's emotion data. As a specific example of operation, the server determines the content of the notification message and sends it to the user's device. The output is a notification of denial or re-examination.
[0678] This ensures that the entire process, from uploading the video file to analysis and notification of the results, proceeds smoothly, and appropriate feedback is provided to the user.
[0679] (Application example 2)
[0680] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0681] In recent years, the distribution of video content has rapidly increased, making the integrity of the content and compliance with legal regulations important issues. However, manual video review is labor-intensive and time-consuming, and has limited scalability. There is also a need for a method to reduce the anxiety and stress users feel when receiving review results. Furthermore, determining whether content is inappropriate or copyright infringing is complex, creating a demand for an accurate and fast automated review system.
[0682] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0683] In this invention, the server includes means for receiving and analyzing video files, means for converting the audio portion of the video file into text data using voice recognition technology, means for analyzing the text data by comparing it with pre-set regulations, means for analyzing the video portion of the video file using image recognition technology, means for determining whether the video is permitted based on the results of the analysis of the text data and video data and issuing a token, means for transmitting the token to a user terminal, means for the user terminal to display the analysis results and provide feedback, and means for recognizing user emotions and assisting the analysis process. This enables automatic review of the soundness of video content and compliance with laws and regulations, and also realizes the provision of feedback that takes user emotions into consideration.
[0684] A "video file" is a file containing video and audio stored in digital format.
[0685] An "analyzing means" is a device or software module for analyzing the contents of a video file and extracting specific information.
[0686] "Speech recognition technology" is a technology for converting voice data into text data.
[0687] "Text data" refers to digital data converted into character information.
[0688] "Regulations" refer to pre-established rules and standards.
[0689] "Image recognition technology" is a technology that analyzes video data and extracts specific information.
[0690] A "token" refers to authentication information issued by the system, which is a digital code that certifies a specific operation or right.
[0691] A "user terminal" is a device such as a computer or smartphone that a user uses to upload video files and receive analysis results.
[0692] A "means for providing feedback" is a device or software module for communicating analysis results or other information to a user.
[0693] The "means for recognizing emotions" is a device or software module for analyzing the user's facial expressions and voice to determine their emotional state.
[0694] The "means for assisting the analysis process" is a device or software module for adjusting the presentation method of the analysis results based on the emotion recognition results.
[0695] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[0696] The automated video review system receives video files uploaded by users and analyzes audio and video data to determine whether they comply with pre-defined regulations. Furthermore, the present invention also incorporates an emotion engine that recognizes user emotions, which assists the analysis process and improves the user experience.
[0697] The server first receives the video file uploaded from the user's device. Next, it converts the audio portion of the video file into text data using speech recognition technology. This speech recognition typically uses existing systems such as Google Speech-to-Text API or IBM Watson Speech to Text. This text data is then analyzed based on set regulations to check, for example, whether it contains inappropriate language or copyrighted content.
[0698] At the same time, the server analyzes the video portion of the video file using image recognition technology, which uses machine learning models such as OpenCV and TensorFlow, to extract specific frames of the video and check for inappropriate content or potential copyright infringement.
[0699] The emotion engine analyzes the user's facial expressions and tone of voice during the video screening to recognize their emotions. Existing services such as Face++ and Microsoft Emotion API are used for emotion recognition. The emotion engine records the user's emotional state when receiving the video screening results and adjusts the way the analysis results are presented.
[0700] If the analysis results show that the video complies with the regulations, the server issues an authentication token to the video. This token certifies that the video is legitimate and can be displayed to viewers by users pasting it into the video description. The user's device receives this token and provides a method for appropriately placing it in the video description.
[0701] On the other hand, if the analysis results in "review" or "denial," a notification is sent to the user's device. The emotion engine recognizes the user's emotions at this time and selects an appropriate notification message to alleviate any discomfort or stress. After receiving the notification, the user can make any necessary corrections and upload the video again for review.
[0702] As a concrete example, consider the case where a user creates a video recording a funny moment during a live broadcast and uploads it to the system. The server analyzes the audio of the video to check for inappropriate language. It also analyzes the video portion to check for any content that may infringe copyright. The emotion engine monitors the user's emotions when receiving the review results and provides appropriate feedback. If there are no problems with the analysis results, the server issues an authentication token for the video and notifies the user. The user can paste this token in the video description to indicate the integrity of the video to viewers.
[0703] Additionally, examples of prompt sentences for the generative AI model to explain the analysis process of this system are shown below.
[0704] Example prompt sentence:
[0705] It analyzes the input mp4 video file and checks the following based on the audio and video data:
[0706] 1. Does it contain inappropriate language (profanity filter)?
[0707] 2. Does it contain copyright infringement (detecting specific images, music, etc.)?
[0708] Once the analysis is complete, determine relevance and generate feedback. Additionally, monitor the user's emotions and provide appropriate feedback.
[0709] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0710] Step 1:
[0711] A user uploads a video file.
[0712] The user selects a video file and uploads it to the system from the device, along with the video file's metadata (e.g., file name, size, format).
[0713] Input: A video file selected by the user
[0714] Output: Uploaded video files and metadata are sent to the server.
[0715] Step 2:
[0716] The server receives the video file and prepares it for analysis.
[0717] The server temporarily stores the received video file and prepares to start the analysis process, specifically separating the audio and video data.
[0718] Input: Uploaded video file
[0719] Output: Audio and video data are separated.
[0720] Step 3:
[0721] The server converts the voice data into text data using voice recognition technology.
[0722] The server inputs the separated voice data into a voice recognition engine (for example, Google Speech-to-Text API) and generates text data from the voice.
[0723] Input: Audio data
[0724] Output: Text data
[0725] Step 4:
[0726] The server analyzes the text data and checks it against regulations.
[0727] The generated text data is compared with the set regulations (filtering inappropriate language, checking copyright infringement, etc.).
[0728] Input: Text data
[0729] Output: Analysis results based on regulations (compliance / non-compliance)
[0730] Step 5:
[0731] The server analyzes the video data using image recognition technology.
[0732] The server feeds the video data into an image recognition engine (e.g., OpenCV or TensorFlow), extracts specific frames, and checks them for inappropriate content or potential copyright infringement.
[0733] Input: Video data
[0734] Output: Analysis results based on image recognition (match / failure)
[0735] Step 6:
[0736] The emotion engine recognizes the user's emotions.
[0737] The server uses an emotion engine to analyze the user's facial expressions and tone of voice and recognize their emotional state before notifying the user of the analysis results.
[0738] Input: User's facial expression data, voice tone data
[0739] Output: User's emotional state data (e.g., relief, tension, anxiety)
[0740] Step 7:
[0741] The server generates feedback based on the analysis results and emotional data.
[0742] The server generates appropriate feedback based on the emotion recognition results and provides the analysis results to the user.
[0743] Input: Analysis results, emotional state data
[0744] Output: Feedback message provided to the user
[0745] Step 8:
[0746] The server determines the video's authorization and issues an authorization token.
[0747] If the video complies with the regulations, the server generates an authentication token and sends it to the user's device.
[0748] Input: All analysis results
[0749] Output: Authentication token
[0750] Step 9:
[0751] The server sends the token to the user terminal.
[0752] Once the authentication token is generated, the server sends it to the user's device, and the user pastes the token into the video description.
[0753] Input: Authentication Token
[0754] Output: Token sent to user device
[0755] Step 10:
[0756] The server sends a notification to the user.
[0757] If the analysis result is "re-examination" or "denial," a notification is sent to the user's device to inform them of the necessary corrections.
[0758] Input: Analysis results (review / rejection)
[0759] Output: Message to be sent to the user
[0760] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0761] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0762] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0763] [Third embodiment]
[0764] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0765] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0766] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0767] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0768] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0769] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0770] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0771] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0772] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0773] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0774] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0775] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0776] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[0777] The automated video review AI system receives video files uploaded by users and analyzes the audio and video data to determine whether they comply with pre-set regulations. The specific program processing is explained below.
[0778] First, the user uploads the edited video file from the device to the platform, and the device then sends the video file to the server, completing the video upload.
[0779] The server receives the uploaded video file and begins analyzing the contents of the video file. Specifically, it converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is obtained as a string of characters.
[0780] The server then analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure it does not contain inappropriate language or copyright-infringing content.
[0781] At the same time, the server uses image recognition technology to analyze the video portion of the video file, extracting specific frames to determine whether the video contains any content that violates regulations, such as copyrighted images or inappropriate content.
[0782] Based on the results of these analyses, the server determines whether the video is permitted. If permitted, the server issues an authentication token for the video. This token certifies that the video complies with the regulations and is stored in a database in association with the video ID.
[0783] The server sends the issued token to the user's device, and the user can paste the token into the description of the video to indicate that the video is appropriate for viewers.
[0784] On the other hand, if the analysis result is "review" or "rejection," the server will send a notification to the user's device. The user can check the notification, make any necessary corrections, and then upload the video again for review.
[0785] As a concrete example, consider the case where a user creates a video by clipping interesting moments from a live broadcast and uploads the video. The server analyzes the audio of the video to check, for example, whether it contains inappropriate language. It also analyzes the video portion to check whether it contains copyright infringement or inappropriate content. If the analysis results are satisfactory, the server issues an authorization token and notifies the user. The user can indicate the integrity of the video by pasting this token in the video's description.
[0786] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations, and supports distributors and platforms in providing sound content.
[0787] The processing flow will be explained below.
[0788] Step 1:
[0789] After the user has finished editing the video, they open the platform's upload screen, click the upload button, select the video file, and click the "Upload" button.
[0790] Step 2:
[0791] The device sends the video file specified by the user to the server. Once the sending is complete, a confirmation message indicating that the upload is complete is displayed on the screen.
[0792] Step 3:
[0793] The server receives the video file sent from the device, checks the video file structure (format and length) and verifies that there are no errors, generates a video ID and stores it in the video database.
[0794] Step 4:
[0795] The server passes the audio portion of the received video file to a speech recognition module to extract the audio data, converts the extracted audio data into text (STT: Speech-to-Text), and temporarily stores the converted text data.
[0796] Step 5:
[0797] The server passes the generated text data to a natural language processing module, which performs profanity filtering and copyright infringement checks, judges the analysis results based on regulations, and records the results.
[0798] Step 6:
[0799] The server passes the video portion of the video file to an image recognition module, which extracts specific frames from the video data and analyzes them for content that violates regulations (e.g., logos, inappropriate content), compares the analysis results with regulations, and records the results.
[0800] Step 7:
[0801] The server combines the results of the text and video analysis to make a comprehensive decision. Based on the decision, if permission is granted, a token is generated. The token is associated with the video ID and stored in a database, and if a re-examination is required, this is recorded.
[0802] Step 8:
[0803] The server sends the review result to the user's device. The result will include either "permit," "review," or "deny." If the request is approved, a URL containing a token will be provided.
[0804] Step 9:
[0805] The terminal receives the examination results sent from the server, displays a notification, and displays a screen where the user can check the examination results.
[0806] Step 10:
[0807] For videos that are allowed, the user pastes the token provided by the server into the summary section of the video, then copies the token into the video description and makes it available to viewers.
[0808] Example 1
[0809] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0810] On online platforms, quickly and accurately determining the appropriateness of video content uploaded by users requires significant resources. Furthermore, the distribution of inappropriate content can undermine the platform's credibility. Traditional methods require manual review, which is labor-intensive and time-consuming. This increases the workload for content review and makes it difficult to distribute content in real time.
[0811] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0812] In this invention, the server includes means for receiving and analyzing video data, means for converting the audio portion of the video data into text data using voice recognition technology, means for analyzing the text data by checking it against preset rules, means for analyzing the video portion of the video data using image recognition technology, means for determining whether the video is permitted based on the results of the analysis of the text data and video data, and means for issuing authentication information and transmitting the authentication information to the user terminal. This allows video content uploaded by users to be quickly and accurately screened, and only appropriate content can be published on the platform.
[0813] "Video data" means media in digital form that contains visual and audio information.
[0814] "Speech recognition technology" is a technology for analyzing voice data and converting it into corresponding text data.
[0815] "Character data" is text-format data generated by voice recognition technology.
[0816] "Regulations" are pre-defined standards or conditions for determining the appropriateness of video content.
[0817] "Image recognition technology" is a technology for analyzing digital images and identifying and recognizing their contents.
[0818] "Authentication information" is digital information issued by a server to prove the authorization of video content.
[0819] "User terminal" means an electronic device used to upload videos and receive authentication information.
[0820] The following describes in detail the embodiments of the present invention. The AI system for automatic video screening receives video data uploaded by users and analyzes the audio and video data to determine whether the data conforms to pre-set regulations. Specific program processing is described below.
[0821] First, the user uploads the video data they have edited from their device to the platform. The device then sends this video data to the server, completing the video upload. The server receives the uploaded video data and begins analyzing its content. Specifically, it converts the audio portion of the video data into text data using voice recognition technology. Through this process, all of the audio information in the video is obtained as a string of characters.
[0822] The server then analyzes the generated text data and verifies whether it complies with established standards. For example, it uses profanity filters and copyright infringement checks to ensure that it does not contain inappropriate language or copyright-infringing content. At the same time, the server uses image recognition technology to analyze the video portion of the video data. This analysis involves extracting specific frames and checking whether the video contains content that violates standards. For example, it checks whether the video contains copyrighted images or inappropriate content.
[0823] Based on these analysis results, the server determines whether the video is permitted. If permission is granted, the server issues authentication information for the video. This authentication information proves that the video complies with the regulations and is stored in a database in association with the video ID. The server then sends the issued authentication information to the user's device. The user can paste the received authentication information into the video's description to indicate to viewers that the video is appropriate. On the other hand, if the analysis result is "review required" or "not permitted," the server sends a notification to the user's device. The user can review the notification, make any necessary corrections, and then upload the video again for review.
[0824] As a concrete example, consider the case where a user creates a video by cutting out interesting moments from a live broadcast and uploads the video. The server analyzes the audio of the video to check for inappropriate language. It also analyzes the video portion to check for copyright infringement or inappropriate content. If the analysis results are satisfactory, the server issues authentication information and notifies the user. The user can then paste this authentication information into the video's description to demonstrate the integrity of the video.
[0825] An example of a prompt for a generative AI model is: "I want to design a system that allows users to upload videos, analyzes the audio and video of the videos, and automatically reviews whether they comply with regulations. The server will perform this analysis using voice recognition and image recognition technology. Please tell me the specific processing steps, the technologies used, and the detailed operations at each step."
[0826] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations, and supports distributors and platforms in providing sound content.
[0827] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0828] Step 1:
[0829] The user uploads the video data after editing from the device to the platform.
[0830] Input: Edited video data
[0831] Output: Notification of completion of upload to the platform
[0832] What it does: After a user completes editing in video editing software (e.g., Adobe Premiere Pro), they open a web browser and use the platform's upload function, which selects and uploads a video file from their local disk.
[0833] Step 2:
[0834] The device sends video data to the server.
[0835] Input: User uploaded video data
[0836] Output: Data transfer to server completed
[0837] Specific operation: When the upload button is clicked, the device sends the video file to the server via an HTTP POST request. An internet connection is required for transmission.
[0838] Step 3:
[0839] The server receives the video data and stores it in storage.
[0840] Input: Video data sent from the device
[0841] Output: Video file saved in storage
[0842] Specific operation: The server saves the received video file to the specified storage (e.g., Amazon S3 bucket). Once saving is complete, metadata is registered in the database.
[0843] Step 4:
[0844] The server converts the audio portion of the video data into text data using voice recognition technology.
[0845] Input: Video file saved in storage
[0846] Output: Text data corresponding to the audio data
[0847] What happens: The server analyzes the video file, extracts the audio, and then converts the audio data into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API).
[0848] Step 5:
[0849] The server parses the generated character data against the rules.
[0850] Input: Text data generated by speech recognition technology
[0851] Output: Judgment result indicating whether the specification is met
[0852] What happens: The server analyzes the text data and performs profanity filters and copyright infringement checks to ensure that it does not contain inappropriate or infringing content.
[0853] Step 6:
[0854] The server analyzes the video portion of the video data using image recognition technology.
[0855] Input: Video file saved in storage
[0856] Output: Analysis results based on video data
[0857] Specific operation: The server extracts specific frames from the video and analyzes the video data using image recognition technology (e.g., OpenCV or Google Cloud Vision API). It checks whether the video contains any illegal content.
[0858] Step 7:
[0859] The server determines whether the video is permitted based on the results of analyzing the text data and video data, and issues authentication information.
[0860] Input: Analysis results of text data and video data after matching
[0861] Output: Certification information if approved, disapproval or reconsideration notice
[0862] Specific operation: The server evaluates the analysis results and determines whether the video complies with the regulations. If it is approved, it generates authentication information and stores it in the database. If it is not approved or requires review, it records that fact.
[0863] Step 8:
[0864] The server sends authentication information or a re-examination or denial notice to the user's device.
[0865] Input: Authorization decision result and authentication information
[0866] Output: Notification to the user's device
[0867] Specific operation: The server sends a notification to the user's device containing authentication information (if authorized) or a notification indicating reconsideration or denial. Notification methods include email and / or in-app notifications.
[0868] Step 9:
[0869] The user reviews the notification they received, makes any necessary corrections, and requests a reconsideration.
[0870] Input: Reconsideration or Denial Notice
[0871] Output: Corrected video data
[0872] What to do: The user will review the notification, make the necessary changes using video editing software, re-upload the video, and request a reconsideration.
[0873] (Application example 1)
[0874] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0875] On conventional video distribution platforms, manually reviewing video audio and video for regulatory compliance required a significant amount of time and effort, making the process highly inefficient. It was also difficult to provide users with prompt feedback, resulting in delays in video release. Furthermore, the platform lacked sufficient guidance and token issuance functions to verify the regulatory compliance of videos uploaded by users.
[0876] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0877] In this invention, the server includes means for receiving and analyzing video files, means for converting the audio portion of the video file into text data using voice recognition technology, means for analyzing the text data by comparing it with preset regulations, means for analyzing the video portion of the video file using image recognition technology, means for determining whether the video is permitted based on the analysis results of the text data and video data and issuing a token, and means for notifying the user terminal of the analysis result and the permission token and providing guidance on re-uploading. This automates video review and enables rapid feedback, allowing users to efficiently upload videos to the video distribution platform and granting tokens to permitted videos.
[0878] A "video file" is a digital file containing video and audio, and is content uploaded by a user.
[0879] "Analysis" refers to the process of individually analyzing the audio and video data of a video file to verify compliance with regulations.
[0880] "Voice recognition technology" is a technology that converts voice data into text data, and is used to obtain voice information within a video as a string of characters.
[0881] "Text data" refers to the text information of the audio data in a video that has been converted using voice recognition technology.
[0882] "Regulations" are rules and standards set in advance by video distribution platforms to evaluate whether videos comply.
[0883] "Verification" is the process of comparing text and video data with regulations to confirm compliance.
[0884] "Image recognition technology" is a technology that analyzes video data and identifies specific objects or scenes.
[0885] "Analysis results" are data that indicates whether the audio and video data conforms to regulations after analysis.
[0886] "Approved" means that the video is deemed to comply with the regulations and is allowed to be made public.
[0887] A "token" is a digital credential that proves a video complies with regulations.
[0888] "Re-uploading" refers to the act of a user re-uploading a video that has been edited to the platform.
[0889] A "guide" is an instruction or advice that tells the user what modifications are needed.
[0890] The present invention relates to an AI system for automatically reviewing videos. This system receives video files uploaded by users and analyzes their audio and video data to determine whether they comply with pre-defined regulations. Specific methods for implementing the system of the present invention are described below.
[0891] First, a user uploads the edited video file to the platform from a device such as a smartphone. The device then sends the video file to the server, completing the video upload.
[0892] The server receives the uploaded video file and begins analyzing the content of the video file. Specifically, it uses the following hardware and software:
[0893] 1. Hardware
[0894] server
[0895] Smartphone (iOS or Android)
[0896] 2. Software
[0897] Python Program
[0898] Audio Speech Recognition (ASR)
[0899] Image recognition technology (computer vision)
[0900] Regulation Check System
[0901] Next, the server converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is captured as a string of characters. For example, the server uses ASR technology to convert speech such as "Today was fun! Please come visit again." into text.
[0902] The server analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure that it does not contain inappropriate language or copyright-infringing content.
[0903] At the same time, the server uses image recognition technology to analyze the video portion of the video file. This analysis involves extracting specific frames to determine whether the video contains any content that violates regulations. For example, it checks for copyrighted images or inappropriate content. The server also performs facial recognition on the video portion to check for the presence of famous characters or specific logos.
[0904] Based on the results of these analyses, the server determines whether the video is permitted. If permission is granted, the server issues an authorization token for the video. This token proves that the video complies with the regulations and is stored in a database in association with the video ID. The server notifies the user of the issued token to their device. Once the user receives it, they can paste the token into the video's description.
[0905] On the other hand, if the analysis result is "review required" or "rejected," the server will send a notification to the user's device. The notification will contain detailed information about the problem and the necessary corrections, so the user can make corrections based on that information. If the user re-uploads the file, it can be reviewed again.
[0906] Example prompt sentence:
[0907] "Extract audio from uploaded videos, transcribe it, and check for specific keywords or phrases."
[0908] "Analyze video frames to see if they contain specific images or logos."
[0909] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations, and supports distributors and platforms in providing sound content.
[0910] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0911] Step 1:
[0912] Users upload the edited video file from their device to the platform.
[0913] Specific operation: A user opens the smartphone app, selects a video file, and taps the upload button, which causes the device to send the video file to the server.
[0914] Input: User's video file
[0915] Output: Video file sent to the server
[0916] Step 2:
[0917] The server receives the uploaded video file and begins analyzing the audio portion of the video file.
[0918] Specific operation: The server extracts the audio data from the video file and converts it into text data using speech recognition technology.
[0919] Input: Video file (audio data)
[0920] Output: Text data of the audio portion
[0921] Step 3:
[0922] The server analyzes the generated text data and checks whether it complies with pre-set regulations.
[0923] Specific operation: The server verifies the text data using a profanity filter and copyright infringement checking system.
[0924] Input: Text data of the audio portion
[0925] Output: Analysis results of text data (match / non-match)
[0926] Step 4:
[0927] The server analyzes the video portion of the video file, extracts specific frames and uses image recognition technology.
[0928] How it works: The server extracts frames from the video and uses computer vision technology to check for facial recognition and the presence of specific logos.
[0929] Input: Video file (video data)
[0930] Output: Analysis results of video data (suitable / unsuitable)
[0931] Step 5:
[0932] The server determines whether the video complies with regulations based on the results of analyzing the audio and video data.
[0933] Specific operation: The server integrates the analysis results and makes a final decision on whether to allow or deny the request.
[0934] Input: Analysis results of audio data, analysis results of video data
[0935] Output: Video approval / disapproval
[0936] Step 6:
[0937] If the video is authorized, the server issues an authorization token for the video.
[0938] What happens: The server generates a unique authentication token for each authorized video and stores it in a database.
[0939] Input: Video permission decision
[0940] Output: Authorization token
[0941] Step 7:
[0942] The server notifies the user terminal of the issued token and provides guidance on how to re-upload.
[0943] Specific operation: The server sends a notification to the user terminal, displays the authorization token along with the analysis result, or provides a correction guide in case of reconsideration or denial.
[0944] Input: Token issuance result
[0945] Output: Notification to user device (token or correction guide)
[0946] Step 8:
[0947] The user can check the notification, make any necessary corrections, and then upload the video again.
[0948] Specific actions: The user will check the notification, make the necessary corrections, and re-upload the video to the platform.
[0949] Input: Notification from the server
[0950] Output: Corrected video file
[0951] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0952] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[0953] The AI video review system receives video files uploaded by users and analyzes audio and video data to determine whether they comply with pre-defined regulations. Furthermore, the invention incorporates an emotion engine that recognizes user emotions, assisting the analysis process and improving the user experience.
[0954] First, the user uploads the edited video file from their device to the platform, and the device sends the video file to the server, displaying a confirmation message on the screen that the upload is complete.
[0955] The server receives the uploaded video file and begins analyzing the contents of the video file. Specifically, it converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is obtained as a string of characters.
[0956] The server then analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure it does not contain inappropriate language or copyright-infringing content.
[0957] At the same time, the server uses image recognition technology to analyze the video portion of the video file, extracting specific frames to determine whether the video contains any content that violates regulations, such as copyrighted images or inappropriate content.
[0958] Furthermore, the system incorporates an emotion engine that analyzes the user's facial expressions and tone of voice to determine their emotions, allowing the system to understand how the analysis results affect the user and provide more appropriate feedback.
[0959] The emotion engine records the user's emotional data and provides feedback based on the analysis. This data is also used as auxiliary information when determining whether a video complies with regulations. For example, if the user is feeling anxious or nervous, the system will carefully convey the analysis results, taking that emotion into consideration.
[0960] Based on the analysis, the server determines whether the video is permitted. If permitted, the server issues an authentication token for the video. This token certifies that the video complies with the regulations and is stored in a database in association with the video ID.
[0961] The server sends the issued token to the user's device, and the user can paste the token into the description of the video to indicate that the video is appropriate for viewers.
[0962] On the other hand, if the analysis result is "review required" or "rejected," a notification is sent to the user's device. The emotion engine recognizes the user's emotions at this time and selects an appropriate notification message to alleviate discomfort or stress, for example. The user can review the notification and make any necessary corrections before uploading the video for review again.
[0963] As a concrete example, consider the case where a user creates a video by cutting out interesting moments from a live broadcast and uploads it. The server analyzes the audio of the video to check for inappropriate language, for example. It also analyzes the video portion to ensure that it does not contain copyright infringement or inappropriate content. The emotion engine monitors the user's emotions when receiving the video review results and provides appropriate feedback. If the analysis results are satisfactory, the server issues an authorization token and notifies the user. The user can paste this token in the video description to indicate the integrity of the video to viewers.
[0964] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations and supports distributors and platforms in providing healthy content. Furthermore, the incorporation of an emotion engine improves the user experience and makes the analysis process run more smoothly.
[0965] The processing flow will be explained below.
[0966] Step 1:
[0967] After the user has finished editing the video, they open the platform's upload screen, click the upload button, select the video file, and click the "Upload" button.
[0968] Step 2:
[0969] The device sends the video file specified by the user to the server. Once the sending is complete, a confirmation message indicating that the upload is complete is displayed on the screen.
[0970] Step 3:
[0971] The server receives the received video file, checks the structure (format and length) of the video file, and verifies that there are no errors. If there are no problems, it generates a video ID and stores it in the video database.
[0972] Step 4:
[0973] The server extracts the audio portion of the video file and converts the audio data into text data using speech recognition technology (STT technology). The converted text data is temporarily stored.
[0974] Step 5:
[0975] The server passes the generated text data to a natural language processing module, which analyzes the text data by checking it against pre-defined regulations, performs profanity filters and copyright infringement checks, and records the analysis results.
[0976] Step 6:
[0977] The server extracts the video portion of the video file and analyzes it using image recognition technology. It extracts specific frames and checks whether they contain any content that violates regulations (e.g., logos, inappropriate content). It then compares the analysis results with regulations and records the results.
[0978] Step 7:
[0979] The server integrates the analysis results and makes a comprehensive decision. Based on this, it determines whether the video is permitted, and if so, generates a token. The token is associated with the video ID and stored in the database.
[0980] Step 8:
[0981] The server uses an emotion engine to analyze the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to determine the user's emotions. The analysis results are recorded and reflected in feedback.
[0982] Step 9:
[0983] The server then takes into account the user's sentiment analysis results and sends a token to the user's device, which contains information indicating that the video complies with the regulations.
[0984] Step 10:
[0985] The terminal receives the audit result and token sent from the server, displays a notification, and displays a screen where the user can check the audit result and token.
[0986] Step 11:
[0987] For videos that are allowed, the user pastes the token provided by the server into the summary section of the video, then copies the token into the video description and makes it available to viewers.
[0988] Step 12:
[0989] If the analysis result is "review" or "rejection," the server sends a notification to the user's device. At this time, the emotion engine monitors the user's emotions and provides appropriate feedback. If necessary, the emotion engine also makes suggestions on how to correct the video. The user checks the notification, makes any necessary corrections, and then uploads the video again for review.
[0990] Example 2
[0991] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0992] With the increase in video content, it is becoming increasingly important to screen videos for inappropriate content and content that infringes copyright. However, manual screening is time-consuming and labor-intensive, and it is often difficult to take appropriate action. In addition, it is necessary to consider the user's feelings regarding the screening results, so a method for performing these tasks automatically is needed.
[0993] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0994] In this invention, the server includes means for receiving and analyzing video files, means for converting the audio portion of the video file into text data using voice recognition technology, means for analyzing the text data by comparing it with pre-set regulations, means for analyzing the video portion of the video file using image recognition technology, means for assisting the analysis process with an emotion engine that recognizes user emotions, means for determining whether the video is permitted based on the analysis results of the text data and video data, and means for issuing an authentication token, and means for transmitting the authentication token to the user terminal. This makes it possible to efficiently check whether the video complies with regulations and improve the user experience.
[0995] Key Word Definitions
[0996] A "video file" is a file format for digital data that contains video and audio.
[0997] "Receiving" refers to the server taking in digital data sent from a user terminal.
[0998] "Analysis" refers to the extraction and evaluation of detailed information from the audio and video data of a video file using specific algorithms and techniques.
[0999] "Speech recognition technology" is a technology for converting voice data into text data.
[1000] "Text data" is data in the form of a string of characters converted using speech recognition technology.
[1001] "Regulations" are pre-established rules and standards for determining the suitability of video content.
[1002] "Verification" is the process of comparing the acquired text and video data with the regulations to see if they match.
[1003] "Image recognition technology" is a technology that uses specific algorithms to identify and analyze objects and scenes in video data.
[1004] An "emotion engine" is a technology that reads emotions from a user's facial expressions, tone of voice, etc., and assists in the analysis process.
[1005] An "authentication token" is electronic proof data issued to prove that a video complies with regulations.
[1006] A "terminal" is an electronic device such as a computer, smartphone, or tablet that is operated by a user.
[1007] "Notification" refers to the act of sending analysis results and other information from the server to the user terminal.
[1008] A "server" is a central processing unit that receives video files, analyzes them, issues authentication tokens, and so on.
[1009] "Feedback" refers to the method and content of the analysis results and notifications provided by the server to the user.
[1010] MODE FOR CARRYING OUT THE INVENTION
[1011] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[1012] System Overview
[1013] The AI video automatic review system receives video files uploaded by users and analyzes audio and video data to determine whether they comply with pre-defined regulations. Furthermore, the present invention also incorporates an emotion engine that recognizes user emotions, assisting the analysis process and improving the user experience.
[1014] Uploading and receiving video files
[1015] First, the user uploads the edited video file from their device to the platform. The device then sends the video file to the server, and a confirmation message appears on the screen confirming the upload. The server then receives the uploaded video file and begins to analyze its contents.
[1016] Audio analysis
[1017] Specifically, the server converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is captured as a string of characters. This process uses speech recognition software such as the Google Cloud Speech-to-Text API.
[1018] Text data regulation check
[1019] The server then analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure it does not contain inappropriate language or copyright-infringing content.
[1020] Video Analysis
[1021] At the same time, the server analyzes the video portion of the video file using image recognition technology. This involves extracting specific frames and identifying whether the video contains any content that violates regulations, such as copyrighted images or inappropriate content. This process is performed using image recognition software such as Amazon Rekognition.
[1022] Emotion engine assists the analysis process
[1023] Furthermore, the system incorporates an emotion engine that analyzes the user's facial expressions and tone of voice to determine their emotions. This allows the system to understand how the analysis results affect the user and provide more appropriate feedback. The emotion engine uses the emotion recognition API of Azure Cognitive Services.
[1024] Notification of analysis results and response
[1025] Based on the analysis results, the server determines whether the video is permitted. If permission is granted, the server issues an authentication token for the video. This token certifies that the video complies with the regulations and is stored in a database in association with the video ID. The server then sends the issued token to the user's device.
[1026] On the other hand, if the analysis result is "review required" or "rejected," a notification is sent to the user's device. The emotion engine recognizes the user's emotions at this time and selects an appropriate notification message to alleviate discomfort or stress. The user can review the notification, make any necessary corrections, and then upload the video again for review.
[1027] Examples of specific examples and prompts
[1028] As a concrete example, consider the case where a user creates a video by cutting out interesting moments from a live broadcast and uploads it. The server analyzes the audio of the video to check for inappropriate language, for example. It also analyzes the video portion to ensure that it does not contain copyright infringement or inappropriate content. The emotion engine monitors the user's emotions when receiving the video review results and provides appropriate feedback. If the analysis results are satisfactory, the server issues an authorization token and notifies the user. The user can paste this token in the video description to indicate the integrity of the video to viewers.
[1029] Prompt Sentence Examples
[1030] "I would like to upload a video file for review. Please make sure that the analysis result does not contain any inappropriate language or copyright infringement. Also, please monitor my emotions when receiving the review result and provide appropriate feedback."
[1031] This system allows users to efficiently check whether videos comply with regulations, enabling broadcasters and platforms to provide healthy content. The built-in emotion engine also improves the user experience and makes the analysis process smoother.
[1032] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1033] Program processing flow
[1034] Step 1: Upload your video file
[1035] The user uploads the edited video file from the device to the platform. As input, there is a video file previously edited by the user, which the device sends to the server. The device transfers the video file to the server and displays the upload progress with a progress bar. As output, the server receives the video file and sends a confirmation message to the device that the upload is complete.
[1036] Step 2: Receiving the video file and preparing it for analysis
[1037] The server receives the uploaded video file and begins preparations for analysis. The input is the video file stored on the server side. Specifically, the server determines where to save the received video file and performs preparations for analysis (storing the file, initializing the necessary analysis modules, etc.). The output is a video file that is ready for analysis.
[1038] Step 3: Audio analysis
[1039] The server extracts the audio portion of the video file and converts the audio data to text using speech recognition technology. The input is the audio data extracted from the video file. For example, the server extracts the audio track using a tool such as ffmpeg and sends the audio data to the Google Cloud Speech-to-Text API. The output is the audio data converted to text.
[1040] Step 4: Regulation check of text data
[1041] The server analyzes the generated text data and checks it against pre-set regulations. The input is text data obtained through speech analysis. Specifically, the server applies a profanity filter to the text data using regular expressions to check for inappropriate words. It then performs text mining to check for the inclusion of words related to copyright. The output is the result of the regulation check.
[1042] Step 5: Video Analysis
[1043] The server extracts the video portion of the video and analyzes the video using image recognition technology. The input is the video data from the video file. For example, the server uses tools such as Amazon Rekognition to extract specific frames and check for violations. The output is the analysis results of the video data.
[1044] Step 6: Sentiment Analysis
[1045] The server uses an emotion engine to analyze the user's emotions. As input, data related to the user's emotions (e.g., facial expressions and tone of voice) is provided. Specifically, the server sends the user's profile picture and video thumbnails to the emotion recognition API to obtain emotion data. As output, the user's emotion data is obtained.
[1046] Step 7: Judging and notifying analysis results
[1047] The server determines whether to allow a video based on the analysis results of the text data and video data. The input is the analysis results of the text data and video data. If permission is granted, the server generates an authentication token, stores it in a database, and sends it to the user's device. The output is an authentication token and a notification. If permission is denied or re-examination is required, the server generates an appropriate notification message and sends a message using an emotion engine that minimizes the impact on the user. The input is the analysis results and the user's emotion data. As a specific example of operation, the server determines the content of the notification message and sends it to the user's device. The output is a notification of denial or re-examination.
[1048] This ensures that the entire process, from uploading the video file to analysis and notification of the results, proceeds smoothly, and appropriate feedback is provided to the user.
[1049] (Application example 2)
[1050] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1051] In recent years, the distribution of video content has rapidly increased, making the integrity of the content and compliance with legal regulations important issues. However, manual video review is labor-intensive and time-consuming, and has limited scalability. There is also a need for a method to reduce the anxiety and stress users feel when receiving review results. Furthermore, determining whether content is inappropriate or copyright infringing is complex, creating a demand for an accurate and fast automated review system.
[1052] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1053] In this invention, the server includes means for receiving and analyzing video files, means for converting the audio portion of the video file into text data using voice recognition technology, means for analyzing the text data by comparing it with pre-set regulations, means for analyzing the video portion of the video file using image recognition technology, means for determining whether the video is permitted based on the results of the analysis of the text data and video data and issuing a token, means for transmitting the token to a user terminal, means for the user terminal to display the analysis results and provide feedback, and means for recognizing user emotions and assisting the analysis process. This enables automatic review of the soundness of video content and compliance with laws and regulations, and also realizes the provision of feedback that takes user emotions into consideration.
[1054] A "video file" is a file containing video and audio stored in digital format.
[1055] An "analyzing means" is a device or software module for analyzing the contents of a video file and extracting specific information.
[1056] "Speech recognition technology" is a technology for converting voice data into text data.
[1057] "Text data" refers to digital data converted into character information.
[1058] "Regulations" refer to pre-established rules and standards.
[1059] "Image recognition technology" is a technology that analyzes video data and extracts specific information.
[1060] A "token" refers to authentication information issued by the system, which is a digital code that certifies a specific operation or right.
[1061] A "user terminal" is a device such as a computer or smartphone that a user uses to upload video files and receive analysis results.
[1062] A "means for providing feedback" is a device or software module for communicating analysis results or other information to a user.
[1063] The "means for recognizing emotions" is a device or software module for analyzing the user's facial expressions and voice to determine their emotional state.
[1064] The "means for assisting the analysis process" is a device or software module for adjusting the presentation method of the analysis results based on the emotion recognition results.
[1065] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[1066] The automated video review system receives video files uploaded by users and analyzes audio and video data to determine whether they comply with pre-defined regulations. Furthermore, the present invention also incorporates an emotion engine that recognizes user emotions, which assists the analysis process and improves the user experience.
[1067] The server first receives the video file uploaded from the user's device. Next, it converts the audio portion of the video file into text data using speech recognition technology. This speech recognition typically uses existing systems such as Google Speech-to-Text API or IBM Watson Speech to Text. This text data is then analyzed based on set regulations to check, for example, whether it contains inappropriate language or copyrighted content.
[1068] At the same time, the server analyzes the video portion of the video file using image recognition technology, which uses machine learning models such as OpenCV and TensorFlow, to extract specific frames of the video and check for inappropriate content or potential copyright infringement.
[1069] The emotion engine analyzes the user's facial expressions and tone of voice during the video screening to recognize their emotions. Existing services such as Face++ and Microsoft Emotion API are used for emotion recognition. The emotion engine records the user's emotional state when receiving the video screening results and adjusts the way the analysis results are presented.
[1070] If the analysis results show that the video complies with the regulations, the server issues an authentication token to the video. This token certifies that the video is legitimate and can be displayed to viewers by users pasting it into the video description. The user's device receives this token and provides a method for appropriately placing it in the video description.
[1071] On the other hand, if the analysis results in "review" or "denial," a notification is sent to the user's device. The emotion engine recognizes the user's emotions at this time and selects an appropriate notification message to alleviate any discomfort or stress. After receiving the notification, the user can make any necessary corrections and upload the video again for review.
[1072] As a concrete example, consider the case where a user creates a video recording a funny moment during a live broadcast and uploads it to the system. The server analyzes the audio of the video to check for inappropriate language. It also analyzes the video portion to check for any content that may infringe copyright. The emotion engine monitors the user's emotions when receiving the review results and provides appropriate feedback. If there are no problems with the analysis results, the server issues an authentication token for the video and notifies the user. The user can paste this token in the video description to indicate the integrity of the video to viewers.
[1073] Additionally, examples of prompt sentences for the generative AI model to explain the analysis process of this system are shown below.
[1074] Example prompt sentence:
[1075] It analyzes the input mp4 video file and checks the following based on the audio and video data:
[1076] 1. Does it contain inappropriate language (profanity filter)?
[1077] 2. Does it contain copyright infringement (detecting specific images, music, etc.)?
[1078] Once the analysis is complete, determine relevance and generate feedback. Additionally, monitor the user's emotions and provide appropriate feedback.
[1079] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1080] Step 1:
[1081] A user uploads a video file.
[1082] The user selects a video file and uploads it to the system from the device, along with the video file's metadata (e.g., file name, size, format).
[1083] Input: A video file selected by the user
[1084] Output: Uploaded video files and metadata are sent to the server.
[1085] Step 2:
[1086] The server receives the video file and prepares it for analysis.
[1087] The server temporarily stores the received video file and prepares to start the analysis process, specifically separating the audio and video data.
[1088] Input: Uploaded video file
[1089] Output: Audio and video data are separated.
[1090] Step 3:
[1091] The server converts the voice data into text data using voice recognition technology.
[1092] The server inputs the separated voice data into a voice recognition engine (for example, Google Speech-to-Text API) and generates text data from the voice.
[1093] Input: Audio data
[1094] Output: Text data
[1095] Step 4:
[1096] The server analyzes the text data and checks it against regulations.
[1097] The generated text data is compared with the set regulations (filtering inappropriate language, checking copyright infringement, etc.).
[1098] Input: Text data
[1099] Output: Analysis results based on regulations (compliance / non-compliance)
[1100] Step 5:
[1101] The server analyzes the video data using image recognition technology.
[1102] The server feeds the video data into an image recognition engine (e.g., OpenCV or TensorFlow), extracts specific frames, and checks them for inappropriate content or potential copyright infringement.
[1103] Input: Video data
[1104] Output: Analysis results based on image recognition (match / failure)
[1105] Step 6:
[1106] The emotion engine recognizes the user's emotions.
[1107] The server uses an emotion engine to analyze the user's facial expressions and tone of voice and recognize their emotional state before notifying the user of the analysis results.
[1108] Input: User's facial expression data, voice tone data
[1109] Output: User's emotional state data (e.g., relief, tension, anxiety)
[1110] Step 7:
[1111] The server generates feedback based on the analysis results and emotional data.
[1112] The server generates appropriate feedback based on the emotion recognition results and provides the analysis results to the user.
[1113] Input: Analysis results, emotional state data
[1114] Output: Feedback message provided to the user
[1115] Step 8:
[1116] The server determines the video's authorization and issues an authorization token.
[1117] If the video complies with the regulations, the server generates an authentication token and sends it to the user's device.
[1118] Input: All analysis results
[1119] Output: Authentication token
[1120] Step 9:
[1121] The server sends the token to the user terminal.
[1122] Once the authentication token is generated, the server sends it to the user's device, and the user pastes the token into the video description.
[1123] Input: Authentication Token
[1124] Output: Token sent to user device
[1125] Step 10:
[1126] The server sends a notification to the user.
[1127] If the analysis result is "re-examination" or "denial," a notification is sent to the user's device to inform them of the necessary corrections.
[1128] Input: Analysis results (review / rejection)
[1129] Output: Message to be sent to the user
[1130] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1131] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1132] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1133] [Fourth embodiment]
[1134] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1135] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1136] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1137] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1138] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1139] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1140] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1141] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1142] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1143] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1144] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1145] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1146] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1147] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[1148] The automated video review AI system receives video files uploaded by users and analyzes the audio and video data to determine whether they comply with pre-set regulations. The specific program processing is explained below.
[1149] First, the user uploads the edited video file from the device to the platform, and the device then sends the video file to the server, completing the video upload.
[1150] The server receives the uploaded video file and begins analyzing the contents of the video file. Specifically, it converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is obtained as a string of characters.
[1151] The server then analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure it does not contain inappropriate language or copyright-infringing content.
[1152] At the same time, the server uses image recognition technology to analyze the video portion of the video file, extracting specific frames to determine whether the video contains any content that violates regulations, such as copyrighted images or inappropriate content.
[1153] Based on the results of these analyses, the server determines whether the video is permitted. If permitted, the server issues an authentication token for the video. This token certifies that the video complies with the regulations and is stored in a database in association with the video ID.
[1154] The server sends the issued token to the user's device, and the user can paste the token into the description of the video to indicate that the video is appropriate for viewers.
[1155] On the other hand, if the analysis result is "review" or "rejection," the server will send a notification to the user's device. The user can check the notification, make any necessary corrections, and then upload the video again for review.
[1156] As a concrete example, consider the case where a user creates a video by clipping interesting moments from a live broadcast and uploads the video. The server analyzes the audio of the video to check, for example, whether it contains inappropriate language. It also analyzes the video portion to check whether it contains copyright infringement or inappropriate content. If the analysis results are satisfactory, the server issues an authorization token and notifies the user. The user can indicate the integrity of the video by pasting this token in the video's description.
[1157] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations, and supports distributors and platforms in providing sound content.
[1158] The processing flow will be explained below.
[1159] Step 1:
[1160] After the user has finished editing the video, they open the platform's upload screen, click the upload button, select the video file, and click the "Upload" button.
[1161] Step 2:
[1162] The device sends the video file specified by the user to the server. Once the sending is complete, a confirmation message indicating that the upload is complete is displayed on the screen.
[1163] Step 3:
[1164] The server receives the video file sent from the device, checks the video file structure (format and length) and verifies that there are no errors, generates a video ID and stores it in the video database.
[1165] Step 4:
[1166] The server passes the audio portion of the received video file to a speech recognition module to extract the audio data, converts the extracted audio data into text (STT: Speech-to-Text), and temporarily stores the converted text data.
[1167] Step 5:
[1168] The server passes the generated text data to a natural language processing module, which performs profanity filtering and copyright infringement checks, judges the analysis results based on regulations, and records the results.
[1169] Step 6:
[1170] The server passes the video portion of the video file to an image recognition module, which extracts specific frames from the video data and analyzes them for content that violates regulations (e.g., logos, inappropriate content), compares the analysis results with regulations, and records the results.
[1171] Step 7:
[1172] The server combines the results of the text and video analysis to make a comprehensive decision. Based on the decision, if permission is granted, a token is generated. The token is associated with the video ID and stored in a database, and if a re-examination is required, this is recorded.
[1173] Step 8:
[1174] The server sends the review result to the user's device. The result will include either "permit," "review," or "deny." If the request is approved, a URL containing a token will be provided.
[1175] Step 9:
[1176] The terminal receives the examination results sent from the server, displays a notification, and displays a screen where the user can check the examination results.
[1177] Step 10:
[1178] For videos that are allowed, the user pastes the token provided by the server into the summary section of the video, then copies the token into the video description and makes it available to viewers.
[1179] Example 1
[1180] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1181] On online platforms, quickly and accurately determining the appropriateness of video content uploaded by users requires significant resources. Furthermore, the distribution of inappropriate content can undermine the platform's credibility. Traditional methods require manual review, which is labor-intensive and time-consuming. This increases the workload for content review and makes it difficult to distribute content in real time.
[1182] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1183] In this invention, the server includes means for receiving and analyzing video data, means for converting the audio portion of the video data into text data using voice recognition technology, means for analyzing the text data by checking it against preset rules, means for analyzing the video portion of the video data using image recognition technology, means for determining whether the video is permitted based on the results of the analysis of the text data and video data, and means for issuing authentication information and transmitting the authentication information to the user terminal. This allows video content uploaded by users to be quickly and accurately screened, and only appropriate content can be published on the platform.
[1184] "Video data" means media in digital form that contains visual and audio information.
[1185] "Speech recognition technology" is a technology for analyzing voice data and converting it into corresponding text data.
[1186] "Character data" is text-format data generated by voice recognition technology.
[1187] "Regulations" are pre-defined standards or conditions for determining the appropriateness of video content.
[1188] "Image recognition technology" is a technology for analyzing digital images and identifying and recognizing their contents.
[1189] "Authentication information" is digital information issued by a server to prove the authorization of video content.
[1190] "User terminal" means an electronic device used to upload videos and receive authentication information.
[1191] The following describes in detail the embodiments of the present invention. The AI system for automatic video screening receives video data uploaded by users and analyzes the audio and video data to determine whether the data conforms to pre-set regulations. Specific program processing is described below.
[1192] First, the user uploads the video data they have edited from their device to the platform. The device then sends this video data to the server, completing the video upload. The server receives the uploaded video data and begins analyzing its content. Specifically, it converts the audio portion of the video data into text data using voice recognition technology. Through this process, all of the audio information in the video is obtained as a string of characters.
[1193] The server then analyzes the generated text data and verifies whether it complies with established standards. For example, it uses profanity filters and copyright infringement checks to ensure that it does not contain inappropriate language or copyright-infringing content. At the same time, the server uses image recognition technology to analyze the video portion of the video data. This analysis involves extracting specific frames and checking whether the video contains content that violates standards. For example, it checks whether the video contains copyrighted images or inappropriate content.
[1194] Based on these analysis results, the server determines whether the video is permitted. If permission is granted, the server issues authentication information for the video. This authentication information proves that the video complies with the regulations and is stored in a database in association with the video ID. The server then sends the issued authentication information to the user's device. The user can paste the received authentication information into the video's description to indicate to viewers that the video is appropriate. On the other hand, if the analysis result is "review required" or "not permitted," the server sends a notification to the user's device. The user can review the notification, make any necessary corrections, and then upload the video again for review.
[1195] As a concrete example, consider the case where a user creates a video by cutting out interesting moments from a live broadcast and uploads the video. The server analyzes the audio of the video to check for inappropriate language. It also analyzes the video portion to check for copyright infringement or inappropriate content. If the analysis results are satisfactory, the server issues authentication information and notifies the user. The user can then paste this authentication information into the video's description to demonstrate the integrity of the video.
[1196] An example of a prompt for a generative AI model is: "I want to design a system that allows users to upload videos, analyzes the audio and video of the videos, and automatically reviews whether they comply with regulations. The server will perform this analysis using voice recognition and image recognition technology. Please tell me the specific processing steps, the technologies used, and the detailed operations at each step."
[1197] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations, and supports distributors and platforms in providing sound content.
[1198] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1199] Step 1:
[1200] The user uploads the video data after editing from the device to the platform.
[1201] Input: Edited video data
[1202] Output: Notification of completion of upload to the platform
[1203] What it does: After a user completes editing in video editing software (e.g., Adobe Premiere Pro), they open a web browser and use the platform's upload function, which selects and uploads a video file from their local disk.
[1204] Step 2:
[1205] The device sends video data to the server.
[1206] Input: User uploaded video data
[1207] Output: Data transfer to server completed
[1208] Specific operation: When the upload button is clicked, the device sends the video file to the server via an HTTP POST request. An internet connection is required for transmission.
[1209] Step 3:
[1210] The server receives the video data and stores it in storage.
[1211] Input: Video data sent from the device
[1212] Output: Video file saved in storage
[1213] Specific operation: The server saves the received video file to the specified storage (e.g., Amazon S3 bucket). Once saving is complete, metadata is registered in the database.
[1214] Step 4:
[1215] The server converts the audio portion of the video data into text data using voice recognition technology.
[1216] Input: Video file saved in storage
[1217] Output: Text data corresponding to the audio data
[1218] What happens: The server analyzes the video file, extracts the audio, and then converts the audio data into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API).
[1219] Step 5:
[1220] The server parses the generated character data against the rules.
[1221] Input: Text data generated by speech recognition technology
[1222] Output: Judgment result indicating whether the specification is met
[1223] What happens: The server analyzes the text data and performs profanity filters and copyright infringement checks to ensure that it does not contain inappropriate or infringing content.
[1224] Step 6:
[1225] The server analyzes the video portion of the video data using image recognition technology.
[1226] Input: Video file saved in storage
[1227] Output: Analysis results based on video data
[1228] Specific operation: The server extracts specific frames from the video and analyzes the video data using image recognition technology (e.g., OpenCV or Google Cloud Vision API). It checks whether the video contains any illegal content.
[1229] Step 7:
[1230] The server determines whether the video is permitted based on the results of analyzing the text data and video data, and issues authentication information.
[1231] Input: Analysis results of text data and video data after matching
[1232] Output: Certification information if approved, disapproval or reconsideration notice
[1233] Specific operation: The server evaluates the analysis results and determines whether the video complies with the regulations. If it is approved, it generates authentication information and stores it in the database. If it is not approved or requires review, it records that fact.
[1234] Step 8:
[1235] The server sends authentication information or a re-examination or denial notice to the user's device.
[1236] Input: Authorization decision result and authentication information
[1237] Output: Notification to the user's device
[1238] Specific operation: The server sends a notification to the user's device containing authentication information (if authorized) or a notification indicating reconsideration or denial. Notification methods include email and / or in-app notifications.
[1239] Step 9:
[1240] The user reviews the notification they received, makes any necessary corrections, and requests a reconsideration.
[1241] Input: Reconsideration or Denial Notice
[1242] Output: Corrected video data
[1243] What to do: The user will review the notification, make the necessary changes using video editing software, re-upload the video, and request a reconsideration.
[1244] (Application example 1)
[1245] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1246] On conventional video distribution platforms, manually reviewing video audio and video for regulatory compliance required a significant amount of time and effort, making the process highly inefficient. It was also difficult to provide users with prompt feedback, resulting in delays in video release. Furthermore, the platform lacked sufficient guidance and token issuance functions to verify the regulatory compliance of videos uploaded by users.
[1247] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1248] In this invention, the server includes means for receiving and analyzing video files, means for converting the audio portion of the video file into text data using voice recognition technology, means for analyzing the text data by comparing it with preset regulations, means for analyzing the video portion of the video file using image recognition technology, means for determining whether the video is permitted based on the analysis results of the text data and video data and issuing a token, and means for notifying the user terminal of the analysis result and the permission token and providing guidance on re-uploading. This automates video review and enables rapid feedback, allowing users to efficiently upload videos to the video distribution platform and granting tokens to permitted videos.
[1249] A "video file" is a digital file containing video and audio, and is content uploaded by a user.
[1250] "Analysis" refers to the process of individually analyzing the audio and video data of a video file to verify compliance with regulations.
[1251] "Voice recognition technology" is a technology that converts voice data into text data, and is used to obtain voice information within a video as a string of characters.
[1252] "Text data" refers to the text information of the audio data in a video that has been converted using voice recognition technology.
[1253] "Regulations" are rules and standards set in advance by video distribution platforms to evaluate whether videos comply.
[1254] "Verification" is the process of comparing text and video data with regulations to confirm compliance.
[1255] "Image recognition technology" is a technology that analyzes video data and identifies specific objects or scenes.
[1256] "Analysis results" are data that indicates whether the audio and video data conforms to regulations after analysis.
[1257] "Approved" means that the video is deemed to comply with the regulations and is allowed to be made public.
[1258] A "token" is a digital credential that proves a video complies with regulations.
[1259] "Re-uploading" refers to the act of a user re-uploading a video that has been edited to the platform.
[1260] A "guide" is an instruction or advice that tells the user what modifications are needed.
[1261] The present invention relates to an AI system for automatically reviewing videos. This system receives video files uploaded by users and analyzes their audio and video data to determine whether they comply with pre-defined regulations. Specific methods for implementing the system of the present invention are described below.
[1262] First, a user uploads the edited video file to the platform from a device such as a smartphone. The device then sends the video file to the server, completing the video upload.
[1263] The server receives the uploaded video file and begins analyzing the content of the video file. Specifically, it uses the following hardware and software:
[1264] 1. Hardware
[1265] server
[1266] Smartphone (iOS or Android)
[1267] 2. Software
[1268] Python Program
[1269] Audio Speech Recognition (ASR)
[1270] Image recognition technology (computer vision)
[1271] Regulation Check System
[1272] Next, the server converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is captured as a string of characters. For example, the server uses ASR technology to convert speech such as "Today was fun! Please come visit again." into text.
[1273] The server analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure that it does not contain inappropriate language or copyright-infringing content.
[1274] At the same time, the server uses image recognition technology to analyze the video portion of the video file. This analysis involves extracting specific frames to determine whether the video contains any content that violates regulations. For example, it checks for copyrighted images or inappropriate content. The server also performs facial recognition on the video portion to check for the presence of famous characters or specific logos.
[1275] Based on the results of these analyses, the server determines whether the video is permitted. If permission is granted, the server issues an authorization token for the video. This token proves that the video complies with the regulations and is stored in a database in association with the video ID. The server notifies the user of the issued token to their device. Once the user receives it, they can paste the token into the video's description.
[1276] On the other hand, if the analysis result is "review required" or "rejected," the server will send a notification to the user's device. The notification will contain detailed information about the problem and the necessary corrections, so the user can make corrections based on that information. If the user re-uploads the file, it can be reviewed again.
[1277] Example prompt sentence:
[1278] "Extract audio from uploaded videos, transcribe it, and check for specific keywords or phrases."
[1279] "Analyze video frames to see if they contain specific images or logos."
[1280] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations, and supports distributors and platforms in providing sound content.
[1281] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1282] Step 1:
[1283] Users upload the edited video file from their device to the platform.
[1284] Specific operation: A user opens the smartphone app, selects a video file, and taps the upload button, which causes the device to send the video file to the server.
[1285] Input: User's video file
[1286] Output: Video file sent to the server
[1287] Step 2:
[1288] The server receives the uploaded video file and begins analyzing the audio portion of the video file.
[1289] Specific operation: The server extracts the audio data from the video file and converts it into text data using speech recognition technology.
[1290] Input: Video file (audio data)
[1291] Output: Text data of the audio portion
[1292] Step 3:
[1293] The server analyzes the generated text data and checks whether it complies with pre-set regulations.
[1294] Specific operation: The server verifies the text data using a profanity filter and copyright infringement checking system.
[1295] Input: Text data of the audio portion
[1296] Output: Analysis results of text data (match / non-match)
[1297] Step 4:
[1298] The server analyzes the video portion of the video file, extracts specific frames and uses image recognition technology.
[1299] How it works: The server extracts frames from the video and uses computer vision technology to check for facial recognition and the presence of specific logos.
[1300] Input: Video file (video data)
[1301] Output: Analysis results of video data (suitable / unsuitable)
[1302] Step 5:
[1303] The server determines whether the video complies with regulations based on the results of analyzing the audio and video data.
[1304] Specific operation: The server integrates the analysis results and makes a final decision on whether to allow or deny the request.
[1305] Input: Analysis results of audio data, analysis results of video data
[1306] Output: Video approval / disapproval
[1307] Step 6:
[1308] If the video is authorized, the server issues an authorization token for the video.
[1309] What happens: The server generates a unique authentication token for each authorized video and stores it in a database.
[1310] Input: Video permission decision
[1311] Output: Authorization token
[1312] Step 7:
[1313] The server notifies the user terminal of the issued token and provides guidance on how to re-upload.
[1314] Specific operation: The server sends a notification to the user terminal, displays the authorization token along with the analysis result, or provides a correction guide in case of reconsideration or denial.
[1315] Input: Token issuance result
[1316] Output: Notification to user device (token or correction guide)
[1317] Step 8:
[1318] The user can check the notification, make any necessary corrections, and then upload the video again.
[1319] Specific actions: The user will check the notification, make the necessary corrections, and re-upload the video to the platform.
[1320] Input: Notification from the server
[1321] Output: Corrected video file
[1322] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1323] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[1324] The AI video review system receives video files uploaded by users and analyzes audio and video data to determine whether they comply with pre-defined regulations. Furthermore, the invention incorporates an emotion engine that recognizes user emotions, assisting the analysis process and improving the user experience.
[1325] First, the user uploads the edited video file from their device to the platform, and the device sends the video file to the server, displaying a confirmation message on the screen that the upload is complete.
[1326] The server receives the uploaded video file and begins analyzing the contents of the video file. Specifically, it converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is obtained as a string of characters.
[1327] The server then analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure it does not contain inappropriate language or copyright-infringing content.
[1328] At the same time, the server uses image recognition technology to analyze the video portion of the video file, extracting specific frames to determine whether the video contains any content that violates regulations, such as copyrighted images or inappropriate content.
[1329] Furthermore, the system incorporates an emotion engine that analyzes the user's facial expressions and tone of voice to determine their emotions, allowing the system to understand how the analysis results affect the user and provide more appropriate feedback.
[1330] The emotion engine records the user's emotional data and provides feedback based on the analysis. This data is also used as auxiliary information when determining whether a video complies with regulations. For example, if the user is feeling anxious or nervous, the system will carefully convey the analysis results, taking that emotion into consideration.
[1331] Based on the analysis, the server determines whether the video is permitted. If permitted, the server issues an authentication token for the video. This token certifies that the video complies with the regulations and is stored in a database in association with the video ID.
[1332] The server sends the issued token to the user's device, and the user can paste the token into the description of the video to indicate that the video is appropriate for viewers.
[1333] On the other hand, if the analysis result is "review required" or "rejected," a notification is sent to the user's device. The emotion engine recognizes the user's emotions at this time and selects an appropriate notification message to alleviate discomfort or stress, for example. The user can review the notification and make any necessary corrections before uploading the video for review again.
[1334] As a concrete example, consider the case where a user creates a video by cutting out interesting moments from a live broadcast and uploads it. The server analyzes the audio of the video to check for inappropriate language, for example. It also analyzes the video portion to ensure that it does not contain copyright infringement or inappropriate content. The emotion engine monitors the user's emotions when receiving the video review results and provides appropriate feedback. If the analysis results are satisfactory, the server issues an authorization token and notifies the user. The user can paste this token in the video description to indicate the integrity of the video to viewers.
[1335] In this way, the system of the present invention allows users to efficiently check whether videos comply with regulations and supports distributors and platforms in providing healthy content. Furthermore, the incorporation of an emotion engine improves the user experience and makes the analysis process run more smoothly.
[1336] The processing flow will be explained below.
[1337] Step 1:
[1338] After the user has finished editing the video, they open the platform's upload screen, click the upload button, select the video file, and click the "Upload" button.
[1339] Step 2:
[1340] The device sends the video file specified by the user to the server. Once the sending is complete, a confirmation message indicating that the upload is complete is displayed on the screen.
[1341] Step 3:
[1342] The server receives the received video file, checks the structure (format and length) of the video file, and verifies that there are no errors. If there are no problems, it generates a video ID and stores it in the video database.
[1343] Step 4:
[1344] The server extracts the audio portion of the video file and converts the audio data into text data using speech recognition technology (STT technology). The converted text data is temporarily stored.
[1345] Step 5:
[1346] The server passes the generated text data to a natural language processing module, which analyzes the text data by checking it against pre-defined regulations, performs profanity filters and copyright infringement checks, and records the analysis results.
[1347] Step 6:
[1348] The server extracts the video portion of the video file and analyzes it using image recognition technology. It extracts specific frames and checks whether they contain any content that violates regulations (e.g., logos, inappropriate content). It then compares the analysis results with regulations and records the results.
[1349] Step 7:
[1350] The server integrates the analysis results and makes a comprehensive decision. Based on this, it determines whether the video is permitted, and if so, generates a token. The token is associated with the video ID and stored in the database.
[1351] Step 8:
[1352] The server uses an emotion engine to analyze the user's emotions. For example, it analyzes the user's facial expressions and tone of voice to determine the user's emotions. The analysis results are recorded and reflected in feedback.
[1353] Step 9:
[1354] The server then takes into account the user's sentiment analysis results and sends a token to the user's device, which contains information indicating that the video complies with the regulations.
[1355] Step 10:
[1356] The terminal receives the audit result and token sent from the server, displays a notification, and displays a screen where the user can check the audit result and token.
[1357] Step 11:
[1358] For videos that are allowed, the user pastes the token provided by the server into the summary section of the video, then copies the token into the video description and makes it available to viewers.
[1359] Step 12:
[1360] If the analysis result is "review" or "rejection," the server sends a notification to the user's device. At this time, the emotion engine monitors the user's emotions and provides appropriate feedback. If necessary, the emotion engine also makes suggestions on how to correct the video. The user checks the notification, makes any necessary corrections, and then uploads the video again for review.
[1361] Example 2
[1362] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1363] With the increase in video content, it is becoming increasingly important to screen videos for inappropriate content and content that infringes copyright. However, manual screening is time-consuming and labor-intensive, and it is often difficult to take appropriate action. In addition, it is necessary to consider the user's feelings regarding the screening results, so a method for performing these tasks automatically is needed.
[1364] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1365] In this invention, the server includes means for receiving and analyzing video files, means for converting the audio portion of the video file into text data using voice recognition technology, means for analyzing the text data by comparing it with pre-set regulations, means for analyzing the video portion of the video file using image recognition technology, means for assisting the analysis process with an emotion engine that recognizes user emotions, means for determining whether the video is permitted based on the analysis results of the text data and video data, and means for issuing an authentication token, and means for transmitting the authentication token to the user terminal. This makes it possible to efficiently check whether the video complies with regulations and improve the user experience.
[1366] Key Word Definitions
[1367] A "video file" is a file format for digital data that contains video and audio.
[1368] "Receiving" refers to the server taking in digital data sent from a user terminal.
[1369] "Analysis" refers to the extraction and evaluation of detailed information from the audio and video data of a video file using specific algorithms and techniques.
[1370] "Speech recognition technology" is a technology for converting voice data into text data.
[1371] "Text data" is data in the form of a string of characters converted using speech recognition technology.
[1372] "Regulations" are pre-established rules and standards for determining the suitability of video content.
[1373] "Verification" is the process of comparing the acquired text and video data with the regulations to see if they match.
[1374] "Image recognition technology" is a technology that uses specific algorithms to identify and analyze objects and scenes in video data.
[1375] An "emotion engine" is a technology that reads emotions from a user's facial expressions, tone of voice, etc., and assists in the analysis process.
[1376] An "authentication token" is electronic proof data issued to prove that a video complies with regulations.
[1377] A "terminal" is an electronic device such as a computer, smartphone, or tablet that is operated by a user.
[1378] "Notification" refers to the act of sending analysis results and other information from the server to the user terminal.
[1379] A "server" is a central processing unit that receives video files, analyzes them, issues authentication tokens, and so on.
[1380] "Feedback" refers to the method and content of the analysis results and notifications provided by the server to the user.
[1381] MODE FOR CARRYING OUT THE INVENTION
[1382] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[1383] System Overview
[1384] The AI video automatic review system receives video files uploaded by users and analyzes audio and video data to determine whether they comply with pre-defined regulations. Furthermore, the present invention also incorporates an emotion engine that recognizes user emotions, assisting the analysis process and improving the user experience.
[1385] Uploading and receiving video files
[1386] First, the user uploads the edited video file from their device to the platform. The device then sends the video file to the server, and a confirmation message appears on the screen confirming the upload. The server then receives the uploaded video file and begins to analyze its contents.
[1387] Audio analysis
[1388] Specifically, the server converts the audio portion of the video file into text data using speech recognition technology. Through this process, all audio information in the video is captured as a string of characters. This process uses speech recognition software such as the Google Cloud Speech-to-Text API.
[1389] Text data regulation check
[1390] The server then analyzes the generated text data and checks whether it complies with established regulations, for example, by using profanity filters and copyright infringement checks to ensure it does not contain inappropriate language or copyright-infringing content.
[1391] Video Analysis
[1392] At the same time, the server analyzes the video portion of the video file using image recognition technology. This involves extracting specific frames and identifying whether the video contains any content that violates regulations, such as copyrighted images or inappropriate content. This process is performed using image recognition software such as Amazon Rekognition.
[1393] Emotion engine assists the analysis process
[1394] Furthermore, the system incorporates an emotion engine that analyzes the user's facial expressions and tone of voice to determine their emotions. This allows the system to understand how the analysis results affect the user and provide more appropriate feedback. The emotion engine uses the emotion recognition API of Azure Cognitive Services.
[1395] Notification of analysis results and response
[1396] Based on the analysis results, the server determines whether the video is permitted. If permission is granted, the server issues an authentication token for the video. This token certifies that the video complies with the regulations and is stored in a database in association with the video ID. The server then sends the issued token to the user's device.
[1397] On the other hand, if the analysis result is "review required" or "rejected," a notification is sent to the user's device. The emotion engine recognizes the user's emotions at this time and selects an appropriate notification message to alleviate discomfort or stress. The user can review the notification, make any necessary corrections, and then upload the video again for review.
[1398] Examples of specific examples and prompts
[1399] As a concrete example, consider the case where a user creates a video by cutting out interesting moments from a live broadcast and uploads it. The server analyzes the audio of the video to check for inappropriate language, for example. It also analyzes the video portion to ensure that it does not contain copyright infringement or inappropriate content. The emotion engine monitors the user's emotions when receiving the video review results and provides appropriate feedback. If the analysis results are satisfactory, the server issues an authorization token and notifies the user. The user can paste this token in the video description to indicate the integrity of the video to viewers.
[1400] Prompt Sentence Examples
[1401] "I would like to upload a video file for review. Please make sure that the analysis result does not contain any inappropriate language or copyright infringement. Also, please monitor my emotions when receiving the review result and provide appropriate feedback."
[1402] This system allows users to efficiently check whether videos comply with regulations, enabling broadcasters and platforms to provide healthy content. The built-in emotion engine also improves the user experience and makes the analysis process smoother.
[1403] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1404] Program processing flow
[1405] Step 1: Upload your video file
[1406] The user uploads the edited video file from the device to the platform. As input, there is a video file previously edited by the user, which the device sends to the server. The device transfers the video file to the server and displays the upload progress with a progress bar. As output, the server receives the video file and sends a confirmation message to the device that the upload is complete.
[1407] Step 2: Receiving the video file and preparing it for analysis
[1408] The server receives the uploaded video file and begins preparations for analysis. The input is the video file stored on the server side. Specifically, the server determines where to save the received video file and performs preparations for analysis (storing the file, initializing the necessary analysis modules, etc.). The output is a video file that is ready for analysis.
[1409] Step 3: Audio analysis
[1410] The server extracts the audio portion of the video file and converts the audio data to text using speech recognition technology. The input is the audio data extracted from the video file. For example, the server extracts the audio track using a tool such as ffmpeg and sends the audio data to the Google Cloud Speech-to-Text API. The output is the audio data converted to text.
[1411] Step 4: Regulation check of text data
[1412] The server analyzes the generated text data and checks it against pre-set regulations. The input is text data obtained through speech analysis. Specifically, the server applies a profanity filter to the text data using regular expressions to check for inappropriate words. It then performs text mining to check for the inclusion of words related to copyright. The output is the result of the regulation check.
[1413] Step 5: Video Analysis
[1414] The server extracts the video portion of the video and analyzes the video using image recognition technology. The input is the video data from the video file. For example, the server uses tools such as Amazon Rekognition to extract specific frames and check for violations. The output is the analysis results of the video data.
[1415] Step 6: Sentiment Analysis
[1416] The server uses an emotion engine to analyze the user's emotions. As input, data related to the user's emotions (e.g., facial expressions and tone of voice) is provided. Specifically, the server sends the user's profile picture and video thumbnails to the emotion recognition API to obtain emotion data. As output, the user's emotion data is obtained.
[1417] Step 7: Judging and notifying analysis results
[1418] The server determines whether to allow a video based on the analysis results of the text data and video data. The input is the analysis results of the text data and video data. If permission is granted, the server generates an authentication token, stores it in a database, and sends it to the user's device. The output is an authentication token and a notification. If permission is denied or re-examination is required, the server generates an appropriate notification message and sends a message using an emotion engine that minimizes the impact on the user. The input is the analysis results and the user's emotion data. As a specific example of operation, the server determines the content of the notification message and sends it to the user's device. The output is a notification of denial or re-examination.
[1419] This ensures that the entire process, from uploading the video file to analysis and notification of the results, proceeds smoothly, and appropriate feedback is provided to the user.
[1420] (Application example 2)
[1421] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1422] In recent years, the distribution of video content has rapidly increased, making the integrity of the content and compliance with legal regulations important issues. However, manual video review is labor-intensive and time-consuming, and has limited scalability. There is also a need for a method to reduce the anxiety and stress users feel when receiving review results. Furthermore, determining whether content is inappropriate or copyright infringing is complex, creating a demand for an accurate and fast automated review system.
[1423] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1424] In this invention, the server includes means for receiving and analyzing video files, means for converting the audio portion of the video file into text data using voice recognition technology, means for analyzing the text data by comparing it with pre-set regulations, means for analyzing the video portion of the video file using image recognition technology, means for determining whether the video is permitted based on the results of the analysis of the text data and video data and issuing a token, means for transmitting the token to a user terminal, means for the user terminal to display the analysis results and provide feedback, and means for recognizing user emotions and assisting the analysis process. This enables automatic review of the soundness of video content and compliance with laws and regulations, and also realizes the provision of feedback that takes user emotions into consideration.
[1425] A "video file" is a file containing video and audio stored in digital format.
[1426] An "analyzing means" is a device or software module for analyzing the contents of a video file and extracting specific information.
[1427] "Speech recognition technology" is a technology for converting voice data into text data.
[1428] "Text data" refers to digital data converted into character information.
[1429] "Regulations" refer to pre-established rules and standards.
[1430] "Image recognition technology" is a technology that analyzes video data and extracts specific information.
[1431] A "token" refers to authentication information issued by the system, which is a digital code that certifies a specific operation or right.
[1432] A "user terminal" is a device such as a computer or smartphone that a user uses to upload video files and receive analysis results.
[1433] A "means for providing feedback" is a device or software module for communicating analysis results or other information to a user.
[1434] The "means for recognizing emotions" is a device or software module for analyzing the user's facial expressions and voice to determine their emotional state.
[1435] The "means for assisting the analysis process" is a device or software module for adjusting the presentation method of the analysis results based on the emotion recognition results.
[1436] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The embodiments for carrying out the present invention will be described in detail below.
[1437] The automated video review system receives video files uploaded by users and analyzes audio and video data to determine whether they comply with pre-defined regulations. Furthermore, the present invention also incorporates an emotion engine that recognizes user emotions, which assists the analysis process and improves the user experience.
[1438] The server first receives the video file uploaded from the user's device. Next, it converts the audio portion of the video file into text data using speech recognition technology. This speech recognition typically uses existing systems such as Google Speech-to-Text API or IBM Watson Speech to Text. This text data is then analyzed based on set regulations to check, for example, whether it contains inappropriate language or copyrighted content.
[1439] At the same time, the server analyzes the video portion of the video file using image recognition technology, which uses machine learning models such as OpenCV and TensorFlow, to extract specific frames of the video and check for inappropriate content or potential copyright infringement.
[1440] The emotion engine analyzes the user's facial expressions and tone of voice during the video screening to recognize their emotions. Existing services such as Face++ and Microsoft Emotion API are used for emotion recognition. The emotion engine records the user's emotional state when receiving the video screening results and adjusts the way the analysis results are presented.
[1441] If the analysis results show that the video complies with the regulations, the server issues an authentication token to the video. This token certifies that the video is legitimate and can be displayed to viewers by users pasting it into the video description. The user's device receives this token and provides a method for appropriately placing it in the video description.
[1442] On the other hand, if the analysis results in "review" or "denial," a notification is sent to the user's device. The emotion engine recognizes the user's emotions at this time and selects an appropriate notification message to alleviate any discomfort or stress. After receiving the notification, the user can make any necessary corrections and upload the video again for review.
[1443] As a concrete example, consider the case where a user creates a video recording a funny moment during a live broadcast and uploads it to the system. The server analyzes the audio of the video to check for inappropriate language. It also analyzes the video portion to check for any content that may infringe copyright. The emotion engine monitors the user's emotions when receiving the review results and provides appropriate feedback. If there are no problems with the analysis results, the server issues an authentication token for the video and notifies the user. The user can paste this token in the video description to indicate the integrity of the video to viewers.
[1444] Additionally, examples of prompt sentences for the generative AI model to explain the analysis process of this system are shown below.
[1445] Example prompt sentence:
[1446] It analyzes the input mp4 video file and checks the following based on the audio and video data:
[1447] 1. Does it contain inappropriate language (profanity filter)?
[1448] 2. Does it contain copyright infringement (detecting specific images, music, etc.)?
[1449] Once the analysis is complete, determine relevance and generate feedback. Additionally, monitor the user's emotions and provide appropriate feedback.
[1450] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1451] Step 1:
[1452] A user uploads a video file.
[1453] The user selects a video file and uploads it to the system from the device, along with the video file's metadata (e.g., file name, size, format).
[1454] Input: A video file selected by the user
[1455] Output: Uploaded video files and metadata are sent to the server.
[1456] Step 2:
[1457] The server receives the video file and prepares it for analysis.
[1458] The server temporarily stores the received video file and prepares to start the analysis process, specifically separating the audio and video data.
[1459] Input: Uploaded video file
[1460] Output: Audio and video data are separated.
[1461] Step 3:
[1462] The server converts the voice data into text data using voice recognition technology.
[1463] The server inputs the separated voice data into a voice recognition engine (for example, Google Speech-to-Text API) and generates text data from the voice.
[1464] Input: Audio data
[1465] Output: Text data
[1466] Step 4:
[1467] The server analyzes the text data and checks it against regulations.
[1468] The generated text data is compared with the set regulations (filtering inappropriate language, checking copyright infringement, etc.).
[1469] Input: Text data
[1470] Output: Analysis results based on regulations (compliance / non-compliance)
[1471] Step 5:
[1472] The server analyzes the video data using image recognition technology.
[1473] The server feeds the video data into an image recognition engine (e.g., OpenCV or TensorFlow), extracts specific frames, and checks them for inappropriate content or potential copyright infringement.
[1474] Input: Video data
[1475] Output: Analysis results based on image recognition (match / failure)
[1476] Step 6:
[1477] The emotion engine recognizes the user's emotions.
[1478] The server uses an emotion engine to analyze the user's facial expressions and tone of voice and recognize their emotional state before notifying the user of the analysis results.
[1479] Input: User's facial expression data, voice tone data
[1480] Output: User's emotional state data (e.g., relief, tension, anxiety)
[1481] Step 7:
[1482] The server generates feedback based on the analysis results and emotional data.
[1483] The server generates appropriate feedback based on the emotion recognition results and provides the analysis results to the user.
[1484] Input: Analysis results, emotional state data
[1485] Output: Feedback message provided to the user
[1486] Step 8:
[1487] The server determines the video's authorization and issues an authorization token.
[1488] If the video complies with the regulations, the server generates an authentication token and sends it to the user's device.
[1489] Input: All analysis results
[1490] Output: Authentication token
[1491] Step 9:
[1492] The server sends the token to the user terminal.
[1493] Once the authentication token is generated, the server sends it to the user's device, and the user pastes the token into the video description.
[1494] Input: Authentication Token
[1495] Output: Token sent to user device
[1496] Step 10:
[1497] The server sends a notification to the user.
[1498] If the analysis result is "re-examination" or "denial," a notification is sent to the user's device to inform them of the necessary corrections.
[1499] Input: Analysis results (review / rejection)
[1500] Output: Message to be sent to the user
[1501] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1502] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1503] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1504] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1505] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1506] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1507] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1508] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1509] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1510] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1511] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1512] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1513] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1514] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1515] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1516] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1517] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1518] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1519] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1520] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1521] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1522] The following is further disclosed regarding the above embodiment.
[1523] (Claim 1)
[1524] means for receiving and analyzing video files;
[1525] means for converting the audio portion of the video file into text data using a voice recognition technology;
[1526] means for analyzing the text data by comparing it with a predetermined regulation;
[1527] means for analyzing the video portion of the video file using image recognition technology;
[1528] A means for determining permission for the video and issuing a token based on the analysis results of the text data and video data;
[1529] means for transmitting the token to a user terminal;
[1530] A system including:
[1531] (Claim 2)
[1532] 10. The system of claim 1, further comprising means for a user to paste the token for an authorized video into a description of the video.
[1533] (Claim 3)
[1534] The system according to claim 1, further comprising means for sending a notification to a user's terminal when the analysis result is "review" or "denial."
[1535] "Example 1"
[1536] (Claim 1)
[1537] means for receiving and analyzing video data;
[1538] means for converting the audio portion of the video data into character data using a voice recognition technique;
[1539] means for analyzing the character data by comparing it with a predetermined rule;
[1540] means for analyzing the video portion of the video data using image recognition technology;
[1541] a means for determining permission for the video based on the analysis results of the character data and the video data and issuing authentication information;
[1542] means for transmitting the authentication information to a user terminal;
[1543] A system including:
[1544] (Claim 2)
[1545] 2. The system according to claim 1, further comprising means for a user to paste the authentication information for an authorized video in a description field of the video.
[1546] (Claim 3)
[1547] The system according to claim 1, further comprising means for sending a notification to a user's terminal when the analysis result is "reexamination" or "denial."
[1548] "Application Example 1"
[1549] Below is the content of the original patent claim with the characteristic parts of the application example added.
[1550] (Claim 1)
[1551] means for receiving and analyzing video files;
[1552] means for converting the audio portion of the video file into text data using a voice recognition technology;
[1553] means for analyzing the text data by comparing it with a predetermined regulation;
[1554] means for analyzing the video portion of the video file using image recognition technology;
[1555] A means for determining permission for the video and issuing a token based on the analysis results of the text data and video data;
[1556] means for notifying a user terminal of the analysis result and the permission token and providing a guide for re-uploading;
[1557] A system including:
[1558] (Claim 2)
[1559] 10. The system of claim 1, further comprising means for a user to paste the token for an authorized video into a description of the video.
[1560] (Claim 3)
[1561] The system according to claim 1, further comprising means for sending a notification to a user's terminal when the analysis result is "review" or "disapproval" and displaying the problem and a correction guide.
[1562] "Example 2: Combining Emotion Engines"
[1563] Rewritten claims
[1564] (Claim 1)
[1565] means for receiving and analyzing video files;
[1566] means for converting the audio portion of the video file into text data using a voice recognition technology;
[1567] means for analyzing the text data by comparing it with a predetermined regulation;
[1568] means for analyzing the video portion of the video file using image recognition technology;
[1569] means for assisting the analysis process with an emotion engine that recognizes the emotions of the user;
[1570] a means for determining permission for the video based on the analysis results of the text data and the video data and issuing an authentication token;
[1571] means for transmitting the authentication token to a user terminal;
[1572] A system including:
[1573] (Claim 2)
[1574] 10. The system of claim 1, further comprising means for a user to paste the authentication token for an authorized video into a description of the video.
[1575] (Claim 3)
[1576] The system according to claim 1, further comprising means for sending a notification to a user's terminal when the analysis result is "review" or "denial."
[1577] (Claim 4)
[1578] 4. The system according to claim 3, further comprising means for generating an appropriate notification message that takes into consideration the user's emotions based on the analysis result of the video file.
[1579] "Application example 2 when combining emotion engines"
[1580] New Claims
[1581] (Claim 1)
[1582] means for receiving and analyzing video files;
[1583] means for converting the audio portion of the video file into text data using a voice recognition technology;
[1584] means for analyzing the text data by comparing it with a predetermined regulation;
[1585] means for analyzing the video portion of the video file using image recognition technology;
[1586] A means for determining permission for the video and issuing a token based on the analysis results of the text data and video data;
[1587] means for transmitting the token to a user terminal;
[1588] means for displaying the analysis results and providing feedback on the user terminal;
[1589] and a means for recognizing a user's emotions and assisting the analysis process.
[1590] (Claim 2)
[1591] 10. The system of claim 1, further comprising means for a user to paste the token for an authorized video into a description of the video.
[1592] (Claim 3)
[1593] The system according to claim 1, further comprising means for sending a notification to a user's terminal when the analysis result is "review" or "denial." [Explanation of symbols]
[1594] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving and analyzing video files; means for converting the audio portion of the video file into text data using a voice recognition technology; means for analyzing the text data by comparing it with a predetermined regulation; means for analyzing the video portion of the video file using image recognition technology; A means for determining permission for the video and issuing a token based on the analysis results of the text data and video data; means for transmitting the token to a user terminal; A system including:
2. 2. The system of claim 1, further comprising means for a user to paste the token for an authorized video into a description of the video.
3. The system according to claim 1 , further comprising means for sending a notification to a user's terminal when the analysis result is "review" or "denial."
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A