system

The system employs AI analysis to evaluate video content for illegality and suggests improvements, addressing the challenge of illegal content circulation by automating detection and response, ensuring efficient and legal compliance.

JP2026070259APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

The rapid proliferation of video content on social networks and platforms leads to the circulation of illegal content due to ambiguous legal interpretations and users' lack of understanding of the law, resulting in privacy violations and other damages, which conventional technologies struggle to detect and respond to effectively.

Method used

A system that uses AI analysis to evaluate the illegality of video data through image and audio recognition, generates an illegality score, and provides users with improvement suggestions, while allowing for pre-detection and controlled publication based on company instructions.

Benefits of technology

Enables efficient and effective management of illegal content by automating the detection and response process, reducing the risk of illegal content being published and providing users with actionable legal improvements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070259000001_ABST
    Figure 2026070259000001_ABST
Patent Text Reader

Abstract

This system automates the pre-detection and response to illegal videos, enabling faster and more efficient management of illegal content. [Solution] A system comprising: means for a user to send video data to a server via a terminal; means for the server to store the video data in storage and prepare it for analysis; means for the server to run the video data through an AI analysis engine and perform image recognition and audio analysis; means for the server to generate a score that evaluates illegality based on the analysis results; means for the server to notify the operating company based on the illegality score; means for the server to send improvement suggestions based on legal grounds to the user; and means for the server to process the video data appropriately based on instructions from the operating company.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern times when a huge amount of video content is posted, there is a problem that illegal videos circulate on SNSs and video platforms. In particular, due to ambiguous legal interpretations and the fact that uploaders do not fully understand the law, there is a high possibility that content with illegality is uploaded regardless of malicious intent. Such content may cause privacy violations and other damages, and thus it has been difficult to appropriately detect and respond to it with conventional technologies.

Means for Solving the Problems

[0005] This invention is a system that allows users to transmit video data via their terminal, which is then stored on a server. The system then uses an AI analysis engine to perform image recognition and audio analysis to pre-evaluate the illegality of the data. Based on the analysis results, the system scores the illegality of the content, notifying the operating company of content that is likely to be illegal, and providing users with legally-based improvement suggestions. Based on instructions from the operating company, the system controls the publication of video data and provides means for re-analyzing and appropriately processing content modified by the user. This system automates the pre-detection and response to illegal videos, enabling faster and more efficient management of illegal content.

[0006] "User" is a term that refers to an individual or group that attempts to submit video data to the system via a device.

[0007] "Terminal" is a term that refers to electronic devices such as computers and smartphones that are used by users to create video data and send it to a server.

[0008] "Server" is a term that refers to a central computer system that receives, analyzes, and processes video data.

[0009] "Video data" is a term that refers to digital files containing moving images and audio recorded by a user using their device.

[0010] "Storage" is a term that refers to a data storage device used by a server to temporarily or permanently hold video data.

[0011] The term "AI analysis engine" refers to a program and its execution environment designed using artificial intelligence technology to analyze video data.

[0012] "Image recognition" is a term that refers to the process of analyzing elements of still images within video data to identify, classify, or detect specific objects.

[0013] "Audio analysis" is a term that refers to the process of analyzing audio contained within video data to identify, classify, or detect its characteristics and content.

[0014] The term "operating company" refers to a corporation or organization that manages and operates the platform on which video data is posted and directs actions against illegal content.

[0015] "Illegality score" is a term that refers to an evaluation index generated based on the results of analysis, which quantifies the likelihood that video data violates laws and regulations.

[0016] "Improvement suggestions" is a term that refers to notifications to users that include recommendations for modifying video data to make it legal content, based on legal grounds. [Brief explanation of the drawing]

[0017] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9]Shows an emotion map to which a plurality of emotions are mapped. [Figure 10] Shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Mode for Carrying Out the Invention

[0018] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.

[0019] First, the language used in the following description will be described.

[0020] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0021] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0025] [First Embodiment]

[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0038] One embodiment of this invention is a system in which a user uses a terminal to send video data to a server, and a series of illegality detection processes are provided based on that data.

[0039] First, the user shoots or edits video data on their device and prepares to post it to the platform. When the user posts, the video data is sent to the server.

[0040] The server stores the received video data in storage and prepares it for the AI ​​analysis engine. The video data is analyzed using image recognition technology to identify elements of still images, and audio analysis technology to identify the content and characteristics of the sound.

[0041] Next, the server uses an AI analysis engine to analyze each frame in the video to check for any elements that violate specific laws or regulations. For example, it determines whether copyrighted music is being used or whether illegal acts or items are depicted.

[0042] Based on the analysis, the server generates a score that evaluates the illegality of the video data. This score indicates the extent to which the video violates existing laws, and if it exceeds a certain threshold, it is deemed illegal.

[0043] Subsequently, the server sends a notification to the operating company regarding the video deemed illegal. The operating company receives this notification and sends instructions to the server regarding whether or not to stop or delete the video.

[0044] Furthermore, the server will also notify users about the illegality and offer suggestions for improvement. These suggestions may include methods for deleting or modifying specific parts of the video.

[0045] Furthermore, based on instructions from the operating company, the server controls the publication of problematic video data, re-analyzes it after user corrections, and allows video posting if appropriate.

[0046] This system allows for the detection and management of illegal video content before it is uploaded, effectively suppressing illegal content on the platform. For example, if illegal use of music is detected in a video uploaded by a user, the system recommends that the offending music be removed, and the upload is permitted only after the removal process is completed.

[0047] The following describes the processing flow.

[0048] Step 1:

[0049] The user uses their device to record or select a video and sends the video data to the server for posting. The video data also includes metadata.

[0050] Step 2:

[0051] The server receives the video data and saves it to storage. This saving is done in preparation for analysis, and data integrity is checked simultaneously.

[0052] Step 3:

[0053] The server passes the video to the AI ​​analysis engine, which then begins image recognition processing. Specifically, it detects specific objects and text from each frame in the video and checks for items that may be legally problematic.

[0054] Step 4:

[0055] The server continues to perform audio analysis using its AI analysis engine. It converts the audio contained in the video into text and detects the use of copyrighted music and illegal words.

[0056] Step 5:

[0057] The server integrates the results of image recognition and audio analysis to generate a score that assesses illegality. This score is calculated based on the importance of each element that violates the law.

[0058] Step 6:

[0059] Based on the illegality score generated by the server, if a certain threshold is exceeded, the operating company will be notified. The notification will include detailed analysis results.

[0060] Step 7:

[0061] The server sends the user a notification based on the analysis results, offering suggestions for improving problematic content or audio. The user then considers correcting the video based on these suggestions.

[0062] Step 8:

[0063] Based on instructions from the operating company, the server will decide to take down, delete, or partially edit the content of the video in question and perform the necessary actions.

[0064] Step 9:

[0065] After the user corrects the video based on the server's notification, the server re-analyzes it. Once it confirms that the problem has been resolved, the server allows the video to be posted or made public.

[0066] (Example 1)

[0067] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0068] Currently, the publication of illegal content in video data posted on online platforms is a serious problem. In particular, videos containing copyright infringement or inappropriate content are often published without prior detection, increasing legal risks for platform users and operators. To solve this problem, it is necessary to efficiently and accurately detect illegality and take appropriate countermeasures.

[0069] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0070] In this invention, the server includes means for a user to transmit video data to an information processing device via an information processing terminal, means for the information processing device to store the video data in a storage device and prepare it for analysis, and means for the information processing device to process the video data through a machine learning analysis device to perform video recognition and audio analysis. This makes it possible to evaluate the possibility of legal violations of the video data before posting it, and to make suggestions or control its publication as needed.

[0071] A "user" is the entity that transmits video data to the system via an information processing terminal.

[0072] An "information processing terminal" is a device used by a user to capture or edit video data and transmit it to an information processing device.

[0073] "Video data" refers to video content submitted by users, and its format and quality must be suitable for analysis on the server.

[0074] An "information processing device" is a computer system that receives transmitted video data and performs analysis and evaluation procedures.

[0075] A "storage device" is a digital storage device used by an information processing device to hold video data received by the device.

[0076] A "machine learning analysis device" is a technological device that performs image recognition and audio analysis on video data to evaluate the possibility of legal violations.

[0077] "Video recognition" is an analytical technology that identifies specific objects or people in a video and detects whether it contains illegal activities or infringing materials.

[0078] "Audio analysis" is a technology that analyzes the audio track within a video to determine whether certain music or sound effects infringe on any rights.

[0079] An "indicator for evaluating the possibility of legal violations" is an evaluation standard that quantifies the extent to which video data violates existing laws and regulations.

[0080] A "corporate entity" refers to an organization that operates the system and issues instructions regarding the publication of video data.

[0081] A "prompt message" is a set of instructions sent through a generative AI model to provide users or companies with legally-based improvement suggestions.

[0082] One embodiment of this invention is a system in which a user transmits video data to an information processing device using an information processing terminal, and based on this data, illegality is detected and managed. Specifically, it is implemented as follows.

[0083] First, the user captures or edits video data using their own information processing terminal. The terminal has video editing software installed, which allows the data to be converted and saved in an appropriate format. The prepared video data is then transmitted to the information processing device via the terminal.

[0084] The information processing device first saves the video data to a storage device. This storage device could be a large-capacity storage system or a cloud-based storage service. Afterward, this data is prepared to be passed to a machine learning analysis device.

[0085] The machine learning analysis system analyzes video data using video recognition and audio analysis technologies. This utilizes existing video analysis software and audio processing libraries. Video recognition identifies specific objects and actions, while audio analysis checks for copyright infringement of music and audio.

[0086] Based on the analysis results, the information processing device generates an index that quantifies the likelihood of legal violations. This index is analyzed by a generating AI model, and an anomaly is reported if it exceeds a certain threshold. The server notifies the company and also sends users legally-based improvement suggestions in the form of prompt messages.

[0087] For example, if copyrighted music is used in a video recorded by a user, the machine learning analysis system will detect this, and the server will suggest to the user that they delete or modify the audio portion. An example of a prompt used in this case is, "Please check if this video is illegal and suggest any parts that need to be corrected."

[0088] This system allows for the effective control of illegal content on the platform by checking the legality of video data before it is released and making timely improvements.

[0089] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0090] Step 1:

[0091] The user captures or edits video data using an information processing terminal.

[0092] Specific operation: The user uses dedicated video editing software to edit the recorded footage with filters and effects.

[0093] Input: Pre-recorded raw video data.

[0094] Data processing: Filtering and applying effects to the video using editing software.

[0095] Output: Video data file ready for submission.

[0096] Step 2:

[0097] The user transmits video data to the information processing device via an information processing terminal.

[0098] Specific operation: The user uses the upload function of the online platform to send video data to the server. Data format conversion is performed as needed.

[0099] Input: Edited video data file.

[0100] Data processing: Data compression and format conversion.

[0101] Output: Video data stored on the server.

[0102] Step 3:

[0103] The server saves the video data to a storage device and prepares it for analysis.

[0104] Specific operation: The server first saves the video data to secure storage and records metadata. It also performs the necessary preprocessing for analysis.

[0105] Input: Video data stored on the server.

[0106] Data processing: Extraction and recording of metadata.

[0107] Output: Stored video data and metadata.

[0108] Step 4:

[0109] The server uses a machine learning analysis system to perform image recognition and audio analysis.

[0110] Specific operation: The server uses video recognition algorithms to analyze objects and actions within the video, and audio analysis techniques to identify and analyze sound sources.

[0111] Input: Video data acquired from the storage device.

[0112] Data processing: Analysis of video frames and feature extraction from audio tracks.

[0113] Output: Analysis results of video and audio.

[0114] Step 5:

[0115] The server generates an index that evaluates the likelihood of legal violations based on the analysis results.

[0116] Specific operation: Based on the analysis results regarding specific legal violations, the server calculates evaluation metrics using a generated AI model and compares them to a threshold.

[0117] Input: Analysis results of video and audio.

[0118] Data processing: Integration of analysis results and calculation of evaluation metrics.

[0119] Output: Evaluation score for legal violations.

[0120] Step 6:

[0121] The server provides improvement suggestions and notifications to users and companies.

[0122] Specific actions: The server generates prompt messages and sends improvement suggestions based on legal grounds to the user. It also notifies the company of illegality.

[0123] Input: Evaluation metric scores and analysis results.

[0124] Data processing: Generating prompt messages and formatting notifications.

[0125] Output: Notifications and suggestions to users and organizations.

[0126] Step 7:

[0127] After user improvements are made, the server will re-analyze the video data and perform appropriate control.

[0128] Specific actions: The improved video data will be analyzed again to check if the problem has been resolved. If there are no problems, publication will be permitted.

[0129] Input: User-modified video data.

[0130] Data processing: Reanalysis and recalculation of evaluation metrics.

[0131] Output: Final evaluation results and decision on whether or not to publish.

[0132] (Application Example 1)

[0133] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0134] In content distribution services, there is a need for a system that can efficiently and reliably detect and address the legality of user-uploaded videos and images before distribution, especially if they may violate legal standards or regulations. However, existing systems primarily focus on post-distribution processing, leaving a risk that distributed content may be published while still containing illegal elements. Therefore, technology is needed to scrutinize the legality of content in advance and automatically take necessary measures.

[0135] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0136] In this invention, the server includes means for a user to transmit a video to a data processing device via an information processing device; means for the data processing device to hold the video in a data storage device and prepare it for analysis; means for the data processing device to run the video through a knowledge processing engine and perform visual recognition and audio analysis; and means for providing recommended alternative materials when a user modifies a video. This enables the legality of a video to be distributed to be detected in advance, and enables the distribution of safe and compliant content.

[0137] A "user" is an entity that transmits moving images using an information processing device and undergoes a process of verifying their legality.

[0138] An "information processing device" refers to a device used by a user to transmit video and image data to a data processing device.

[0139] "Motion images" refer to a series of images and audio data that include visual and auditory elements.

[0140] A "data processing unit" is a central system that receives transmitted video and performs various processes for analysis and storage.

[0141] A "data storage device" is a storage medium used to temporarily or permanently store moving images.

[0142] A "knowledge processing engine" is an algorithmic system that automatically analyzes the visual and auditory elements of moving images to determine their legality.

[0143] "Visual recognition" refers to the process of analyzing individual frames within a video or image to identify elements and patterns within that image.

[0144] "Audio analysis" is the process of analyzing audio data contained in video and determining its content and characteristics.

[0145] "Lawfulness" refers to the state in which a video or image does not violate any current laws or regulations.

[0146] A "numerical value" is a numerical evaluation indicating legality, and is an indicator of how problematic the video is in light of legal standards.

[0147] The "operating organization" is the organization responsible for managing and operating the content distribution service.

[0148] "Alternative material" refers to visual or auditory elements that can be used in place of the original video footage, as proposed to ensure legality.

[0149] The system for realizing this invention operates using a program with the following configuration.

[0150] The server receives video footage transmitted by the user using an information processing device. This video footage is temporarily stored in a data storage device. The server then transfers the stored video footage to a knowledge processing engine for visual recognition and audio analysis. This involves using multiple APIs that combine image and audio analysis technologies, with specific examples including Amazon Rekognition and Google Cloud Video Intelligence, depending on the use case.

[0151] Visual recognition analyzes video frames to identify image elements. Audio analysis analyzes audio tracks to identify their content. The server evaluates these analysis results and generates a numerical value indicating legality. This value is compared to current laws and regulations to determine how legal the video is.

[0152] If a user uploads a video containing illegal elements, the server will send the user a legally-based suggestion for improvement. This suggestion will include providing alternative material and showing how to correct or replace the problematic parts.

[0153] For example, if a video shot by a user on their smartphone contains copyrighted music, the server will recognize the music and suggest alternative music that can be used. A prompt message such as "This video contains copyrighted music. Please suggest alternative music that can be used" will be displayed to the user through the system, allowing the user to resolve legal issues and safely publish the video.

[0154] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0155] Step 1:

[0156] The user prepares video footage using an information processing device. The input here is video footage captured or edited by the user. To send this data to the data processing device, the video footage is encoded into an optimal format and prepared for transmission.

[0157] Step 2:

[0158] The terminal transmits video data to the data processing unit. The input is encoded video data, and the output is temporary storage in a data storage device by the data processing unit. Data is transmitted via a communication line, received by the data processing unit, and stored in the data storage device.

[0159] Step 3:

[0160] The server passes the video footage stored in the data storage device to the knowledge processing engine. The input is the stored video footage, and the output is data labeling in preparation for analysis. At this stage, the frames of the video footage are identified and each is converted into an analyzable format.

[0161] Step 4:

[0162] The server uses a knowledge processing engine to perform visual recognition and audio analysis of video. Inputs are labeled frames and audio data, and output is the analysis results based on each element. Visual recognition techniques identify objects and scenes within frames, while audio analysis techniques convert audio tracks into text and features. APIs such as Amazon Rekognition and Google Cloud Video Intelligence are used.

[0163] Step 5:

[0164] The server evaluates the legality of the video based on the analysis results. The input is the analysis results from step 4, and the output is a numerical score indicating the legality of the video. The score is compared to a pre-set legal standard to identify the frames and audio elements in which violations were detected.

[0165] Step 6:

[0166] The server sends the user legally-based improvement suggestions. The input is a legality score and a list of elements that need improvement, and the output is a notification and improvement suggestions for the user. Specifically, a prompt message is generated stating, "This video contains copyrighted music. Please recommend usable alternative music."

[0167] Step 7:

[0168] The user resends the improved video image to the server. The input is the corrected video image data, and the output is a notification that it is ready for re-analysis. The user corrects the video image based on the server's suggestions and resends it to complete the compliance verification process.

[0169] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0170] In this embodiment of the invention, a system is provided that, in addition to pre-evaluating video data, also has a function to analyze the user's emotions. The user uses a terminal to shoot or select a video and sends the video data to the server. Here, the server uses an emotion engine in addition to conventional analysis to analyze the user's emotions.

[0171] The server saves the video data to storage and prepares it for transfer to the AI ​​analysis engine. First, it performs image recognition and audio analysis to check for standard legal violations. Next, it uses an emotion engine to analyze how the video content and audio reflect the user's emotions. For example, it determines whether the user is expressing emotions such as anger, sadness, or joy based on their facial expressions and tone of voice in the video.

[0172] Once the analysis is complete, the server generates an illegality score and performs an overall evaluation that also takes into account additional sentiment analysis results. The sentiment analysis results contribute to determining whether the video content is intentionally illegal and provide additional information for notifications to the operating company and users.

[0173] The server then notifies the operating company of the illegality and sentiment analysis, and instructs them to take down or delete the video as necessary. The server also sends improvement suggestions to the user based on the analysis results, along with legal grounds and emotional feedback. This helps users not only correct illegality but also understand how their emotions are influencing the content. For example, if a user films a video while excited and it contains highly provocative content, the user will receive a message suggesting calm and constructive ways to improve the video.

[0174] The following describes the processing flow.

[0175] Step 1:

[0176] The user uses their device to record or select a video and sends the video data to the server. The user then enters basic information for posting and presses the submit button.

[0177] Step 2:

[0178] The server receives the video data and saves it to storage. During the saving process, the video metadata is extracted to prepare it for analysis.

[0179] Step 3:

[0180] The server activates the AI ​​analysis engine and first performs image recognition processing. This identifies potentially illegal elements from each frame in the video.

[0181] Step 4:

[0182] The server performs audio analysis and converts the audio in the video into text. This allows for the detection of potentially copyrighted music and illegal statements.

[0183] Step 5:

[0184] The server uses an emotion engine to analyze the user's emotions expressed in the video. It analyzes facial expressions, tone of voice, and other factors to determine the user's emotional state.

[0185] Step 6:

[0186] The server integrates the results of image recognition, audio analysis, and sentiment analysis to generate a video illegality score. In addition, it adjusts the overall score to reflect the sentiment analysis results.

[0187] Step 7:

[0188] The server will send a detailed notification to the operating company based on the illegality score and sentiment analysis results. Based on the information included, the operating company will consider whether or not to publish the video.

[0189] Step 8:

[0190] The server sends a notification of the analysis results to the user and provides suggestions for improvement. These suggestions include identifying illegal activities and providing feedback on sentiment, prompting the user to make appropriate corrections.

[0191] Step 9:

[0192] The server receives instructions from the operating company and, as necessary, takes the video data down, deletes it, or prepares it for re-uploading after correction. Videos corrected by the user are re-analyzed, and if the problem is resolved, re-uploading is permitted.

[0193] (Example 2)

[0194] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0195] In modern society, the rapid spread of video content has increased the risk of illegal or emotionally inappropriate content being published. This puts a strain on the resources of regulatory organizations and makes it difficult for users to make informed judgments about the legality and emotional impact of the content they are viewing.

[0196] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0197] In this invention, the server includes means for a user to transmit video information to a computer via a terminal, means for the computer to store the video information in a storage device and prepare it for analysis, and means for the computer to process the video information through an artificial intelligence analysis device to perform image recognition and audio analysis. This makes it possible to comprehensively analyze the illegality and emotional impact of video content, reduce the burden on management organizations, and provide useful feedback to users.

[0198] A "user" is an individual or group that manipulates video information via a terminal and transmits it to a computer.

[0199] A "terminal" is a device used by a user to capture or select video information.

[0200] "Video information" refers to video data that a user sends from their device to a computer.

[0201] A "computer" is a digital device that processes received video information and prepares it for analysis.

[0202] A "memory device" is a digital storage medium used to hold image information within a computer.

[0203] An "artificial intelligence analysis device" is a system that uses video information to perform image recognition and audio analysis.

[0204] "Image recognition" is the process of recognizing visual elements within video information and analyzing the data related to them.

[0205] "Audio analysis" is the process of analyzing the audio within video information and evaluating its content.

[0206] An "evaluation value" is a numerical value or indicator used to assess illegality based on the analysis results.

[0207] A "management organization" is an organization responsible for the appropriate handling and disclosure of video information.

[0208] A "notification" is information or an alert sent from a computer to an administrative organization or user.

[0209] "Improvement suggestions" are guidelines provided to users regarding specific changes or modifications to the content of video information.

[0210] "Feedback" refers to evaluations and opinions provided to users based on analysis results.

[0211] "Initial filtering" is the first process of selecting video information to be analyzed based on thresholds.

[0212] "Emotional analysis" is the process of analyzing a user's emotional state using visual and auditory elements within video information.

[0213] A "threshold" is a numerical value used as a criterion for selecting the target for analysis during the initial filtering process.

[0214] "Publication control" is the process of managing whether or not video information can be provided to viewers.

[0215] "Re-analysis" is the process of re-analyzing video information after the user has made improvements and evaluating the results.

[0216] This system begins with the user using a terminal to capture or select video information and sending that data to a server. The terminal includes camera functions and a file browser to assist with capturing and selecting video information.

[0217] The server stores the received video information in a storage device and prepares it for transfer to the artificial intelligence analysis device. Here, the storage device uses cloud storage or local server storage to securely hold the digital data. The server processes the video information through the artificial intelligence analysis device, which sequentially performs image recognition and audio analysis. Computer vision technology is used for image recognition to recognize objects and backgrounds in the video, while natural language processing technology is used for audio analysis to convert audio data into text and evaluate its content.

[0218] Furthermore, based on the analysis results, the server uses an emotion analysis engine to evaluate the user's emotional state. The emotion analysis engine incorporates facial recognition technology and voice tone analysis capabilities to analyze emotional elements within the video information in detail.

[0219] Once the analysis is complete, the server generates an evaluation score to assess the illegality and notifies the management organization of the results. At the same time, it sends users legally-based improvement suggestions and feedback that takes into account emotional impact, prompting them to take necessary action.

[0220] For example, if a user-created video content is deemed excessively offensive, the server will provide a suggestion for improvement such as, "This video contains offensive content. Please revise it to use milder language." Another example of a prompt for the generative AI model might be, "Analyze this video to determine its illegality or emotional impact, and generate appropriate feedback."

[0221] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0222] Step 1:

[0223] The user captures or selects video information using the device. The device acquires video information through its camera function and provides the ability to select the video file chosen by the user from local storage. It handles the video information selected by the user as input and generates data in a format to be sent to the server as output.

[0224] Step 2:

[0225] The terminal uploads the video information selected by the user to the server. During this process, it verifies that the video information is in the correct format and performs format conversion if necessary. The input is raw data stored on the terminal, and the output is digital data formatted for the server to receive.

[0226] Step 3:

[0227] The server stores the received video information in a storage device and prepares it for analysis. The server saves the files to secure cloud storage or local storage. The input is digital data transmitted from the terminal, and the output is a video file securely stored in the storage device.

[0228] Step 4:

[0229] The server inputs video information stored in its memory into an artificial intelligence analysis device, which then performs image recognition and audio analysis. Image recognition extracts visual elements, while audio analysis converts audio data into text and analyzes its content. The input is digital data including video and audio, and the output consists of identified image objects and analyzed audio text.

[0230] Step 5:

[0231] The server evaluates the user's emotional state using an emotion analysis engine based on the analyzed results. This involves using facial recognition technology to infer emotions from facial expressions in the video and analyzing the emotional elements of the voice using voice tone analysis. The input is the result data from image recognition and voice analysis, and the output is evaluation data regarding the user's emotional state.

[0232] Step 6:

[0233] The server generates an illegality rating based on the analysis results and sentiment evaluation, and sends a notification to the management organization. The rating is used as an indicator of the appropriateness of the video's publication. The input is all the analysis data, and the output is the notification to the management organization.

[0234] Step 7:

[0235] The server sends users legally-based improvement suggestions and provides feedback that takes emotional impact into consideration. The specific suggestions are customized based on the analysis results. Inputs are evaluation values ​​and analysis data, and output is a feedback message to the user.

[0236] (Application Example 2)

[0237] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0238] In modern information and communication technology, user-generated video content often contains inappropriate material, including legal violations, or may be deemed inappropriate due to the influence of user emotions. In such cases, there is a need for a system that provides appropriate filtering before video publication and constructive feedback to users. However, existing technologies lack sufficient means to analyze user emotions and conduct a comprehensive evaluation in conjunction with legal violations. Therefore, there is a need to provide a system that analyzes the legal and emotional aspects of video data and proactively controls and improves the transmission of inappropriate content.

[0239] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0240] In this invention, the server includes means for a user to transmit video data to an information processing device via an information processing device, means for the information processing device to analyze the user's facial expressions and voice and determine their emotions, and means for the information processing device to evaluate the influence of the user's emotional state on the content of the video data. This makes it possible to detect in advance if the video data is affected by legal violations or inappropriate emotions, and to provide appropriate feedback.

[0241] An "information processing device" is a computer system used for receiving, storing, analyzing, and notifying the results of data.

[0242] A "memory device" is a medium used by an information processing device to temporarily or permanently store video data.

[0243] A "data analysis device" is a system used by an information processing device to analyze video data, specifically for analyzing images and audio.

[0244] An "evaluation indicator" is a numerical value or indicator used for evaluation, generated based on legal compliance or other criteria.

[0245] An "organization" is an organization or group that receives notifications from information processing devices and issues instructions regarding the publication or processing of video data.

[0246] "User" refers to an individual or entity that transmits video data to an information processing device and receives analysis and feedback.

[0247] "Emotional discrimination" is the process of analyzing a user's facial expressions and voice to identify their emotional state.

[0248] "Initial filtering" refers to the process by which an information processing device performs basic screening before analyzing video data.

[0249] To implement this invention, a system centered on an information processing device is constructed. Its main components are as follows:

[0250] First, the user captures or selects video data using a smartphone or other device and sends it to the information processing device. This information processing device incorporates a storage device, a data analysis device, and the necessary software. Once the video data is received, the information processing device saves the data to its storage device and prepares it for analysis.

[0251] The data analysis device is equipped with an AI analysis engine for image recognition and speech analysis. Specifically, it uses the DeepFace library to analyze facial expressions and the SpeechRecognition library to analyze the content of speech data. Based on this data, the information processing device determines the user's emotional state and further evaluates how emotions influence the content of the video data.

[0252] Next, the information processing device generates evaluation indicators for legal violations and provides appropriate feedback to the user based on the evaluation results. This includes a function to judge the appropriateness of the video data based on the evaluation indicators and to provide specific improvement suggestions if improvements are needed.

[0253] For example, if a user films an event while excited, the system automatically determines whether the background audio and facial expressions are offensive. If offensive elements are detected, the information processing device sends feedback to the user, such as "We recommend editing this video to make it more positive," to encourage improvement.

[0254] Examples of prompts for generative AI models include the following:

[0255] "A user has requested an emotion analysis of the following video. Please analyze the facial expressions and tone of voice in the video to identify the main emotions. If the emotions are negative, please provide specific advice on how to correct them."

[0256] In this way, it becomes possible to provide comprehensive evaluation and improvement support for video data via information processing equipment.

[0257] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0258] Step 1:

[0259] The user uses a device to capture or select video data. The input is video data from the user's device. The device sends this data to the information processing device. The output is the video data sent to the server.

[0260] Step 2:

[0261] The server saves the received video data to its storage device. The input is video data transmitted from the terminal, which the server temporarily stores in its storage device. The output is data ready for analysis.

[0262] Step 3:

[0263] The server passes the video data stored in its memory to the AI ​​analysis engine, initiating image recognition and audio analysis. The input is the stored video data. The server then uses the DeepFace library to analyze facial expressions and the SpeechRecognition library to convert audio data into text. The output consists of the user's facial expression data and the content of the audio.

[0264] Step 4:

[0265] The server determines the user's emotions based on the analysis results. The input consists of image recognition and voice analysis results, which the server uses to estimate the user's emotions. The output provides the user's main emotions (e.g., joy, anger, sadness).

[0266] Step 5:

[0267] The server evaluates the legal compliance and emotional state of the video data and generates an evaluation index. Inputs include a checklist of legal violations and the results of the emotional assessment. Based on this, the server generates an evaluation index that quantifies the safety and appropriateness of the video data. The output is the evaluation index.

[0268] Step 6:

[0269] The server provides feedback to the user based on evaluation metrics it generates. Inputs include evaluation metrics and improvement suggestion templates. The server generates prompt messages, creates specific feedback such as "We recommend editing this video to be more positive," and sends it to the user. Output is the feedback message sent to the user.

[0270] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0271] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0272] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0273] [Second Embodiment]

[0274] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0275] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0276] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0277] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0278] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0279] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0280] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0281] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0282] The specific processing program 56 is an example of the "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by operating as the specific processing unit 290 according to the specific processing program 56 executed by the processor 28 on the RAM 30.

[0283] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the specific processing unit 290.

[0284] In the smart glasses 214, the receiving and output processing is performed by the processor 46. The storage 50 stores a receiving and output program 60. The processor 46 reads the receiving and output program 60 from the storage 50 and executes the read receiving and output program 60 on the RAM 48. The receiving and output processing is realized by operating as the control unit 46A according to the receiving and output program 60 executed by the processor 46 on the RAM 48.

[0285] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".

[0286] An embodiment of this invention is a system that provides a series of illegality detection processes based on a user transmitting video data to a server using a terminal.

[0287] First, the user shoots or edits video data with their own terminal and prepares to post it on the platform. When the user executes the post, the video data is transmitted to the server.

[0288] The server stores the received video data in storage and prepares it for the AI ​​analysis engine. The video data is analyzed using image recognition technology to identify elements of still images, and audio analysis technology to identify the content and characteristics of the sound.

[0289] Next, the server uses an AI analysis engine to analyze each frame in the video to check for any elements that violate specific laws or regulations. For example, it determines whether copyrighted music is being used or whether illegal acts or items are depicted.

[0290] Based on the analysis, the server generates a score that evaluates the illegality of the video data. This score indicates the extent to which the video violates existing laws, and if it exceeds a certain threshold, it is deemed illegal.

[0291] Subsequently, the server sends a notification to the operating company regarding the video deemed illegal. The operating company receives this notification and sends instructions to the server regarding whether or not to stop or delete the video.

[0292] Furthermore, the server will also notify users about the illegality and offer suggestions for improvement. These suggestions may include methods for deleting or modifying specific parts of the video.

[0293] Furthermore, based on instructions from the operating company, the server controls the publication of problematic video data, re-analyzes it after user corrections, and allows video posting if appropriate.

[0294] This system allows for the detection and management of illegal video content before it is uploaded, effectively suppressing illegal content on the platform. For example, if illegal use of music is detected in a video uploaded by a user, the system recommends that the offending music be removed, and the upload is permitted only after the removal process is completed.

[0295] The following describes the processing flow.

[0296] Step 1:

[0297] The user uses the terminal to shoot or select a video and sends the video data for posting to the server. The video data also includes metadata.

[0298] Step 2:

[0299] The server receives the video data and stores it in storage. The storage is performed in preparation for analysis, and data integrity is also confirmed simultaneously.

[0300] Step 3:

[0301] The server passes the video to the AI analysis engine and starts the image recognition process. Specifically, specific objects and texts are detected from each frame in the video, and items that may pose legal problems are checked.

[0302] Step 4:

[0303] The server continues to perform voice analysis using the AI analysis engine. The voice included in the video is texturized, and the use of copyrighted music and illegal words is detected.

[0304] Step 5:

[0305] The server integrates the results of image recognition and voice analysis and generates a score for evaluating illegality. This score is calculated based on the importance of each legal violation element.

[0306] Step 6:

[0307] Based on the illegality score generated by the server, if it exceeds a certain threshold, the operating company is notified. The notice includes details of the analysis results.

[0308] Step 7:

[0309] The server sends the user a notification based on the analysis results, offering suggestions for improving problematic content or audio. The user then considers correcting the video based on these suggestions.

[0310] Step 8:

[0311] Based on instructions from the operating company, the server will decide to take down, delete, or partially edit the content of the video in question and perform the necessary actions.

[0312] Step 9:

[0313] After the user corrects the video based on the server's notification, the server re-analyzes it. Once it confirms that the problem has been resolved, the server allows the video to be posted or made public.

[0314] (Example 1)

[0315] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0316] Currently, the publication of illegal content in video data posted on online platforms is a serious problem. In particular, videos containing copyright infringement or inappropriate content are often published without prior detection, increasing legal risks for platform users and operators. To solve this problem, it is necessary to efficiently and accurately detect illegality and take appropriate countermeasures.

[0317] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0318] In this invention, the server includes means for a user to transmit video data to an information processing device via an information processing terminal, means for the information processing device to store the video data in a storage device and prepare it for analysis, and means for the information processing device to process the video data through a machine learning analysis device to perform video recognition and audio analysis. This makes it possible to evaluate the possibility of legal violations of the video data before posting it, and to make suggestions or control its publication as needed.

[0319] A "user" is the entity that transmits video data to the system via an information processing terminal.

[0320] An "information processing terminal" is a device used by a user to capture or edit video data and transmit it to an information processing device.

[0321] "Video data" refers to video content submitted by users, and its format and quality must be suitable for analysis on the server.

[0322] An "information processing device" is a computer system that receives transmitted video data and performs analysis and evaluation procedures.

[0323] A "storage device" is a digital storage device used by an information processing device to hold video data received by the device.

[0324] A "machine learning analysis device" is a technological device that performs image recognition and audio analysis on video data to evaluate the possibility of legal violations.

[0325] "Video recognition" is an analytical technology that identifies specific objects or people in a video and detects whether it contains illegal activities or infringing materials.

[0326] "Audio analysis" is a technology that analyzes the audio track within a video to determine whether certain music or sound effects infringe on any rights.

[0327] An "indicator for evaluating the possibility of legal violations" is an evaluation standard that quantifies the extent to which video data violates existing laws and regulations.

[0328] A "corporate entity" refers to an organization that operates the system and issues instructions regarding the publication of video data.

[0329] A "prompt message" is a set of instructions sent through a generative AI model to provide users or companies with legally-based improvement suggestions.

[0330] One embodiment of this invention is a system in which a user transmits video data to an information processing device using an information processing terminal, and based on this data, illegality is detected and managed. Specifically, it is implemented as follows.

[0331] First, the user captures or edits video data using their own information processing terminal. The terminal has video editing software installed, which allows the data to be converted and saved in an appropriate format. The prepared video data is then transmitted to the information processing device via the terminal.

[0332] The information processing device first saves the video data to a storage device. This storage device could be a large-capacity storage system or a cloud-based storage service. Afterward, this data is prepared to be passed to a machine learning analysis device.

[0333] The machine learning analysis system analyzes video data using video recognition and audio analysis technologies. This utilizes existing video analysis software and audio processing libraries. Video recognition identifies specific objects and actions, while audio analysis checks for copyright infringement of music and audio.

[0334] Based on the analysis results, the information processing device generates an index that quantifies the likelihood of legal violations. This index is analyzed by a generating AI model, and an anomaly is reported if it exceeds a certain threshold. The server notifies the company and also sends users legally-based improvement suggestions in the form of prompt messages.

[0335] For example, if copyrighted music is used in a video recorded by a user, the machine learning analysis system will detect this, and the server will suggest to the user that they delete or modify the audio portion. An example of a prompt used in this case is, "Please check if this video is illegal and suggest any parts that need to be corrected."

[0336] This system allows for the effective control of illegal content on the platform by checking the legality of video data before it is released and making timely improvements.

[0337] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0338] Step 1:

[0339] The user captures or edits video data using an information processing terminal.

[0340] Specific operation: The user uses dedicated video editing software to edit the recorded footage with filters and effects.

[0341] Input: Pre-recorded raw video data.

[0342] Data processing: Filtering and applying effects to the video using editing software.

[0343] Output: Video data file ready for submission.

[0344] Step 2:

[0345] The user transmits video data to the information processing device via an information processing terminal.

[0346] Specific operation: The user uses the upload function of the online platform to send video data to the server. Data format conversion is performed as needed.

[0347] Input: Edited video data file.

[0348] Data processing: Data compression and format conversion.

[0349] Output: Video data stored on the server.

[0350] Step 3:

[0351] The server saves the video data to a storage device and prepares it for analysis.

[0352] Specific operation: The server first saves the video data to secure storage and records metadata. It also performs the necessary preprocessing for analysis.

[0353] Input: Video data stored on the server.

[0354] Data processing: Extraction and recording of metadata.

[0355] Output: Stored video data and metadata.

[0356] Step 4:

[0357] The server uses a machine learning analysis system to perform image recognition and audio analysis.

[0358] Specific operation: The server uses video recognition algorithms to analyze objects and actions within the video, and audio analysis techniques to identify and analyze sound sources.

[0359] Input: Video data acquired from the storage device.

[0360] Data processing: Analysis of video frames and feature extraction from audio tracks.

[0361] Output: Analysis results of video and audio.

[0362] Step 5:

[0363] The server generates an index that evaluates the likelihood of legal violations based on the analysis results.

[0364] Specific operation: Based on the analysis results regarding specific legal violations, the server calculates evaluation metrics using a generated AI model and compares them to a threshold.

[0365] Input: Analysis results of video and audio.

[0366] Data processing: Integration of analysis results and calculation of evaluation metrics.

[0367] Output: Evaluation score for legal violations.

[0368] Step 6:

[0369] The server provides improvement suggestions and notifications to users and companies.

[0370] Specific actions: The server generates prompt messages and sends improvement suggestions based on legal grounds to the user. It also notifies the company of illegality.

[0371] Input: Evaluation metric scores and analysis results.

[0372] Data processing: Generating prompt messages and formatting notifications.

[0373] Output: Notifications and suggestions to users and organizations.

[0374] Step 7:

[0375] After user improvements are made, the server will re-analyze the video data and perform appropriate control.

[0376] Specific actions: The improved video data will be analyzed again to check if the problem has been resolved. If there are no problems, publication will be permitted.

[0377] Input: User-modified video data.

[0378] Data processing: Reanalysis and recalculation of evaluation metrics.

[0379] Output: Final evaluation results and decision on whether or not to publish.

[0380] (Application Example 1)

[0381] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0382] In content distribution services, there is a need for a system that can efficiently and reliably detect and address the legality of user-uploaded videos and images before distribution, especially if they may violate legal standards or regulations. However, existing systems primarily focus on post-distribution processing, leaving a risk that distributed content may be published while still containing illegal elements. Therefore, technology is needed to scrutinize the legality of content in advance and automatically take necessary measures.

[0383] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0384] In this invention, the server includes means for a user to transmit a video to a data processing device via an information processing device; means for the data processing device to hold the video in a data storage device and prepare it for analysis; means for the data processing device to run the video through a knowledge processing engine and perform visual recognition and audio analysis; and means for providing recommended alternative materials when a user modifies a video. This enables the legality of a video to be distributed to be detected in advance, and enables the distribution of safe and compliant content.

[0385] A "user" is an entity that transmits moving images using an information processing device and undergoes a process of verifying their legality.

[0386] An "information processing device" refers to a device used by a user to transmit video and image data to a data processing device.

[0387] "Motion images" refer to a series of images and audio data that include visual and auditory elements.

[0388] A "data processing unit" is a central system that receives transmitted video and performs various processes for analysis and storage.

[0389] A "data storage device" is a storage medium used to temporarily or permanently store moving images.

[0390] A "knowledge processing engine" is an algorithmic system that automatically analyzes the visual and auditory elements of moving images to determine their legality.

[0391] "Visual recognition" refers to the process of analyzing individual frames within a video or image to identify elements and patterns within that image.

[0392] "Audio analysis" is the process of analyzing audio data contained in video and determining its content and characteristics.

[0393] "Lawfulness" refers to the state in which a video or image does not violate any current laws or regulations.

[0394] A "numerical value" is a numerical evaluation indicating legality, and is an indicator of how problematic the video is in light of legal standards.

[0395] The "operating organization" is the organization responsible for managing and operating the content distribution service.

[0396] "Alternative material" refers to visual or auditory elements that can be used in place of the original video footage, as proposed to ensure legality.

[0397] The system for realizing this invention operates using a program with the following configuration.

[0398] The server receives video footage transmitted by the user using an information processing device. This video footage is temporarily stored in a data storage device. The server then transfers the stored video footage to a knowledge processing engine for visual recognition and audio analysis. This involves using multiple APIs that combine image and audio analysis technologies, with specific examples including Amazon Rekognition and Google Cloud Video Intelligence, depending on the use case.

[0399] Visual recognition analyzes video frames to identify image elements. Audio analysis analyzes audio tracks to identify their content. The server evaluates these analysis results and generates a numerical value indicating legality. This value is compared to current laws and regulations to determine how legal the video is.

[0400] If a user uploads a video containing illegal elements, the server will send the user a legally-based suggestion for improvement. This suggestion will include providing alternative material and showing how to correct or replace the problematic parts.

[0401] For example, if a video shot by a user on their smartphone contains copyrighted music, the server will recognize the music and suggest alternative music that can be used. A prompt message such as "This video contains copyrighted music. Please suggest alternative music that can be used" will be displayed to the user through the system, allowing the user to resolve legal issues and safely publish the video.

[0402] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0403] Step 1:

[0404] The user prepares video footage using an information processing device. The input here is video footage captured or edited by the user. To send this data to the data processing device, the video footage is encoded into an optimal format and prepared for transmission.

[0405] Step 2:

[0406] The terminal transmits video data to the data processing unit. The input is encoded video data, and the output is temporary storage in a data storage device by the data processing unit. Data is transmitted via a communication line, received by the data processing unit, and stored in the data storage device.

[0407] Step 3:

[0408] The server passes the video footage stored in the data storage device to the knowledge processing engine. The input is the stored video footage, and the output is data labeling in preparation for analysis. At this stage, the frames of the video footage are identified and each is converted into an analyzable format.

[0409] Step 4:

[0410] The server uses a knowledge processing engine to perform visual recognition and audio analysis of video. Inputs are labeled frames and audio data, and output is the analysis results based on each element. Visual recognition techniques identify objects and scenes within frames, while audio analysis techniques convert audio tracks into text and features. APIs such as Amazon Rekognition and Google Cloud Video Intelligence are used.

[0411] Step 5:

[0412] The server evaluates the legality of the video based on the analysis results. The input is the analysis results from step 4, and the output is a numerical score indicating the legality of the video. The score is compared to a pre-set legal standard to identify the frames and audio elements in which violations were detected.

[0413] Step 6:

[0414] The server sends the user legally-based improvement suggestions. The input is a legality score and a list of elements that need improvement, and the output is a notification and improvement suggestions for the user. Specifically, a prompt message is generated stating, "This video contains copyrighted music. Please recommend usable alternative music."

[0415] Step 7:

[0416] The user resends the improved video image to the server. The input is the corrected video image data, and the output is a notification that it is ready for re-analysis. The user corrects the video image based on the server's suggestions and resends it to complete the compliance verification process.

[0417] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0418] In this embodiment of the invention, a system is provided that, in addition to pre-evaluating video data, also has a function to analyze the user's emotions. The user uses a terminal to shoot or select a video and sends the video data to the server. Here, the server uses an emotion engine in addition to conventional analysis to analyze the user's emotions.

[0419] The server saves the video data to storage and prepares it for transfer to the AI ​​analysis engine. First, it performs image recognition and audio analysis to check for standard legal violations. Next, it uses an emotion engine to analyze how the video content and audio reflect the user's emotions. For example, it determines whether the user is expressing emotions such as anger, sadness, or joy based on their facial expressions and tone of voice in the video.

[0420] Once the analysis is complete, the server generates an illegality score and performs an overall evaluation that also takes into account additional sentiment analysis results. The sentiment analysis results contribute to determining whether the video content is intentionally illegal and provide additional information for notifications to the operating company and users.

[0421] The server then notifies the operating company of the illegality and sentiment analysis, and instructs them to take down or delete the video as necessary. The server also sends improvement suggestions to the user based on the analysis results, along with legal grounds and emotional feedback. This helps users not only correct illegality but also understand how their emotions are influencing the content. For example, if a user films a video while excited and it contains highly provocative content, the user will receive a message suggesting calm and constructive ways to improve the video.

[0422] The following describes the processing flow.

[0423] Step 1:

[0424] The user uses their device to record or select a video and sends the video data to the server. The user then enters basic information for posting and presses the submit button.

[0425] Step 2:

[0426] The server receives the video data and saves it to storage. During the saving process, the video metadata is extracted to prepare it for analysis.

[0427] Step 3:

[0428] The server activates the AI ​​analysis engine and first performs image recognition processing. This identifies potentially illegal elements from each frame in the video.

[0429] Step 4:

[0430] The server performs audio analysis and converts the audio in the video into text. This allows for the detection of potentially copyrighted music and illegal statements.

[0431] Step 5:

[0432] The server uses an emotion engine to analyze the user's emotions expressed in the video. It analyzes facial expressions, tone of voice, and other factors to determine the user's emotional state.

[0433] Step 6:

[0434] The server integrates the results of image recognition, audio analysis, and sentiment analysis to generate a video illegality score. In addition, it adjusts the overall score to reflect the sentiment analysis results.

[0435] Step 7:

[0436] The server will send a detailed notification to the operating company based on the illegality score and sentiment analysis results. Based on the information included, the operating company will consider whether or not to publish the video.

[0437] Step 8:

[0438] The server sends a notification of the analysis results to the user and provides suggestions for improvement. These suggestions include identifying illegal activities and providing feedback on sentiment, prompting the user to make appropriate corrections.

[0439] Step 9:

[0440] The server receives instructions from the operating company and, as necessary, takes the video data down, deletes it, or prepares it for re-uploading after correction. Videos corrected by the user are re-analyzed, and if the problem is resolved, re-uploading is permitted.

[0441] (Example 2)

[0442] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0443] In modern society, the rapid spread of video content has increased the risk of illegal or emotionally inappropriate content being published. This puts a strain on the resources of regulatory organizations and makes it difficult for users to make informed judgments about the legality and emotional impact of the content they are viewing.

[0444] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0445] In this invention, the server includes means for a user to transmit video information to a computer via a terminal, means for the computer to store the video information in a storage device and prepare it for analysis, and means for the computer to process the video information through an artificial intelligence analysis device to perform image recognition and audio analysis. This makes it possible to comprehensively analyze the illegality and emotional impact of video content, reduce the burden on management organizations, and provide useful feedback to users.

[0446] A "user" is an individual or group that manipulates video information via a terminal and transmits it to a computer.

[0447] A "terminal" is a device used by a user to capture or select video information.

[0448] "Video information" refers to video data that a user sends from their device to a computer.

[0449] A "computer" is a digital device that processes received video information and prepares it for analysis.

[0450] A "memory device" is a digital storage medium used to hold image information within a computer.

[0451] An "artificial intelligence analysis device" is a system that uses video information to perform image recognition and audio analysis.

[0452] "Image recognition" is the process of recognizing visual elements within video information and analyzing the data related to them.

[0453] "Audio analysis" is the process of analyzing the audio within video information and evaluating its content.

[0454] An "evaluation value" is a numerical value or indicator used to assess illegality based on the analysis results.

[0455] A "management organization" is an organization responsible for the appropriate handling and disclosure of video information.

[0456] A "notification" is information or an alert sent from a computer to an administrative organization or user.

[0457] "Improvement suggestions" are guidelines provided to users regarding specific changes or modifications to the content of video information.

[0458] "Feedback" refers to evaluations and opinions provided to users based on analysis results.

[0459] "Initial filtering" is the first process of selecting video information to be analyzed based on thresholds.

[0460] "Emotional analysis" is the process of analyzing a user's emotional state using visual and auditory elements within video information.

[0461] A "threshold" is a numerical value used as a criterion for selecting the target for analysis during the initial filtering process.

[0462] "Publication control" is the process of managing whether or not video information can be provided to viewers.

[0463] "Re-analysis" is the process of re-analyzing video information after the user has made improvements and evaluating the results.

[0464] This system begins with the user using a terminal to capture or select video information and sending that data to a server. The terminal includes camera functions and a file browser to assist with capturing and selecting video information.

[0465] The server stores the received video information in a storage device and prepares it for transfer to the artificial intelligence analysis device. Here, the storage device uses cloud storage or local server storage to securely hold the digital data. The server processes the video information through the artificial intelligence analysis device, which sequentially performs image recognition and audio analysis. Computer vision technology is used for image recognition to recognize objects and backgrounds in the video, while natural language processing technology is used for audio analysis to convert audio data into text and evaluate its content.

[0466] Furthermore, based on the analysis results, the server uses an emotion analysis engine to evaluate the user's emotional state. The emotion analysis engine incorporates facial recognition technology and voice tone analysis capabilities to analyze emotional elements within the video information in detail.

[0467] Once the analysis is complete, the server generates an evaluation score to assess the illegality and notifies the management organization of the results. At the same time, it sends users legally-based improvement suggestions and feedback that takes into account emotional impact, prompting them to take necessary action.

[0468] For example, if a user-created video content is deemed excessively offensive, the server will provide a suggestion for improvement such as, "This video contains offensive content. Please revise it to use milder language." Another example of a prompt for the generative AI model might be, "Analyze this video to determine its illegality or emotional impact, and generate appropriate feedback."

[0469] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0470] Step 1:

[0471] The user captures or selects video information using the device. The device acquires video information through its camera function and provides the ability to select the video file chosen by the user from local storage. It handles the video information selected by the user as input and generates data in a format to be sent to the server as output.

[0472] Step 2:

[0473] The terminal uploads the video information selected by the user to the server. During this process, it verifies that the video information is in the correct format and performs format conversion if necessary. The input is raw data stored on the terminal, and the output is digital data formatted for the server to receive.

[0474] Step 3:

[0475] The server stores the received video information in a storage device and prepares it for analysis. The server saves the files to secure cloud storage or local storage. The input is digital data transmitted from the terminal, and the output is a video file securely stored in the storage device.

[0476] Step 4:

[0477] The server inputs video information stored in its memory into an artificial intelligence analysis device, which then performs image recognition and audio analysis. Image recognition extracts visual elements, while audio analysis converts audio data into text and analyzes its content. The input is digital data including video and audio, and the output consists of identified image objects and analyzed audio text.

[0478] Step 5:

[0479] The server evaluates the user's emotional state using an emotion analysis engine based on the analyzed results. This involves using facial recognition technology to infer emotions from facial expressions in the video and analyzing the emotional elements of the voice using voice tone analysis. The input is the result data from image recognition and voice analysis, and the output is evaluation data regarding the user's emotional state.

[0480] Step 6:

[0481] The server generates an illegality rating based on the analysis results and sentiment evaluation, and sends a notification to the management organization. The rating is used as an indicator of the appropriateness of the video's publication. The input is all the analysis data, and the output is the notification to the management organization.

[0482] Step 7:

[0483] The server sends users legally-based improvement suggestions and provides feedback that takes emotional impact into consideration. The specific suggestions are customized based on the analysis results. Inputs are evaluation values ​​and analysis data, and output is a feedback message to the user.

[0484] (Application Example 2)

[0485] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0486] In modern information and communication technology, user-generated video content often contains inappropriate material, including legal violations, or may be deemed inappropriate due to the influence of user emotions. In such cases, there is a need for a system that provides appropriate filtering before video publication and constructive feedback to users. However, existing technologies lack sufficient means to analyze user emotions and conduct a comprehensive evaluation in conjunction with legal violations. Therefore, there is a need to provide a system that analyzes the legal and emotional aspects of video data and proactively controls and improves the transmission of inappropriate content.

[0487] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0488] In this invention, the server includes means for a user to transmit video data to an information processing device via an information processing device, means for the information processing device to analyze the user's facial expressions and voice and determine their emotions, and means for the information processing device to evaluate the influence of the user's emotional state on the content of the video data. This makes it possible to detect in advance if the video data is affected by legal violations or inappropriate emotions, and to provide appropriate feedback.

[0489] An "information processing device" is a computer system used for receiving, storing, analyzing, and notifying the results of data.

[0490] A "memory device" is a medium used by an information processing device to temporarily or permanently store video data.

[0491] A "data analysis device" is a system used by an information processing device to analyze video data, specifically for analyzing images and audio.

[0492] An "evaluation indicator" is a numerical value or indicator used for evaluation, generated based on legal compliance or other criteria.

[0493] An "organization" is an organization or group that receives notifications from information processing devices and issues instructions regarding the publication or processing of video data.

[0494] "User" refers to an individual or entity that transmits video data to an information processing device and receives analysis and feedback.

[0495] "Emotional discrimination" is the process of analyzing a user's facial expressions and voice to identify their emotional state.

[0496] "Initial filtering" refers to the process by which an information processing device performs basic screening before analyzing video data.

[0497] To implement this invention, a system centered on an information processing device is constructed. Its main components are as follows:

[0498] First, the user captures or selects video data using a smartphone or other device and sends it to the information processing device. This information processing device incorporates a storage device, a data analysis device, and the necessary software. Once the video data is received, the information processing device saves the data to its storage device and prepares it for analysis.

[0499] The data analysis device is equipped with an AI analysis engine for image recognition and speech analysis. Specifically, it uses the DeepFace library to analyze facial expressions and the SpeechRecognition library to analyze the content of speech data. Based on this data, the information processing device determines the user's emotional state and further evaluates how emotions influence the content of the video data.

[0500] Next, the information processing device generates evaluation indicators for legal violations and provides appropriate feedback to the user based on the evaluation results. This includes a function to judge the appropriateness of the video data based on the evaluation indicators and to provide specific improvement suggestions if improvements are needed.

[0501] For example, if a user films an event while excited, the system automatically determines whether the background audio and facial expressions are offensive. If offensive elements are detected, the information processing device sends feedback to the user, such as "We recommend editing this video to make it more positive," to encourage improvement.

[0502] Examples of prompts for generative AI models include the following:

[0503] "A user has requested an emotion analysis of the following video. Please analyze the facial expressions and tone of voice in the video to identify the main emotions. If the emotions are negative, please provide specific advice on how to correct them."

[0504] In this way, it becomes possible to provide comprehensive evaluation and improvement support for video data via information processing equipment.

[0505] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0506] Step 1:

[0507] The user uses a device to capture or select video data. The input is video data from the user's device. The device sends this data to the information processing device. The output is the video data sent to the server.

[0508] Step 2:

[0509] The server saves the received video data to its storage device. The input is video data transmitted from the terminal, which the server temporarily stores in its storage device. The output is data ready for analysis.

[0510] Step 3:

[0511] The server passes the video data stored in its memory to the AI ​​analysis engine, initiating image recognition and audio analysis. The input is the stored video data. The server then uses the DeepFace library to analyze facial expressions and the SpeechRecognition library to convert audio data into text. The output consists of the user's facial expression data and the content of the audio.

[0512] Step 4:

[0513] The server determines the user's emotions based on the analysis results. The input consists of image recognition and voice analysis results, which the server uses to estimate the user's emotions. The output provides the user's main emotions (e.g., joy, anger, sadness).

[0514] Step 5:

[0515] The server evaluates the legal compliance and emotional state of the video data and generates an evaluation index. Inputs include a checklist of legal violations and the results of the emotional assessment. Based on this, the server generates an evaluation index that quantifies the safety and appropriateness of the video data. The output is the evaluation index.

[0516] Step 6:

[0517] The server provides feedback to the user based on evaluation metrics it generates. Inputs include evaluation metrics and improvement suggestion templates. The server generates prompt messages, creates specific feedback such as "We recommend editing this video to be more positive," and sends it to the user. Output is the feedback message sent to the user.

[0518] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0519] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0520] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0521] [Third Embodiment]

[0522] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0523] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0524] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0525] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0526] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0527] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0528] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0529] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0530] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0531] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0532] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0533] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0534] One embodiment of this invention is a system in which a user uses a terminal to send video data to a server, and a series of illegality detection processes are provided based on that data.

[0535] First, the user shoots or edits video data on their device and prepares to post it to the platform. When the user posts, the video data is sent to the server.

[0536] The server stores the received video data in storage and prepares it for the AI ​​analysis engine. The video data is analyzed using image recognition technology to identify elements of still images, and audio analysis technology to identify the content and characteristics of the sound.

[0537] Next, the server uses an AI analysis engine to analyze each frame in the video to check for any elements that violate specific laws or regulations. For example, it determines whether copyrighted music is being used or whether illegal acts or items are depicted.

[0538] Based on the analysis, the server generates a score that evaluates the illegality of the video data. This score indicates the extent to which the video violates existing laws, and if it exceeds a certain threshold, it is deemed illegal.

[0539] Subsequently, the server sends a notification to the operating company regarding the video deemed illegal. The operating company receives this notification and sends instructions to the server regarding whether or not to stop or delete the video.

[0540] Furthermore, the server will also notify users about the illegality and offer suggestions for improvement. These suggestions may include methods for deleting or modifying specific parts of the video.

[0541] Furthermore, based on instructions from the operating company, the server controls the publication of problematic video data, re-analyzes it after user corrections, and allows video posting if appropriate.

[0542] This system allows for the detection and management of illegal video content before it is uploaded, effectively suppressing illegal content on the platform. For example, if illegal use of music is detected in a video uploaded by a user, the system recommends that the offending music be removed, and the upload is permitted only after the removal process is completed.

[0543] The following describes the processing flow.

[0544] Step 1:

[0545] The user uses their device to record or select a video and sends the video data to the server for posting. The video data also includes metadata.

[0546] Step 2:

[0547] The server receives the video data and saves it to storage. This saving is done in preparation for analysis, and data integrity is checked simultaneously.

[0548] Step 3:

[0549] The server passes the video to the AI ​​analysis engine, which then begins image recognition processing. Specifically, it detects specific objects and text from each frame in the video and checks for items that may be legally problematic.

[0550] Step 4:

[0551] The server continues to perform audio analysis using its AI analysis engine. It converts the audio contained in the video into text and detects the use of copyrighted music and illegal words.

[0552] Step 5:

[0553] The server integrates the results of image recognition and audio analysis to generate a score that assesses illegality. This score is calculated based on the importance of each element that violates the law.

[0554] Step 6:

[0555] Based on the illegality score generated by the server, if a certain threshold is exceeded, the operating company will be notified. The notification will include detailed analysis results.

[0556] Step 7:

[0557] The server sends the user a notification based on the analysis results, offering suggestions for improving problematic content or audio. The user then considers correcting the video based on these suggestions.

[0558] Step 8:

[0559] Based on instructions from the operating company, the server will decide to take down, delete, or partially edit the content of the video in question and perform the necessary actions.

[0560] Step 9:

[0561] After the user corrects the video based on the server's notification, the server re-analyzes it. Once it confirms that the problem has been resolved, the server allows the video to be posted or made public.

[0562] (Example 1)

[0563] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0564] Currently, the publication of illegal content in video data posted on online platforms is a serious problem. In particular, videos containing copyright infringement or inappropriate content are often published without prior detection, increasing legal risks for platform users and operators. To solve this problem, it is necessary to efficiently and accurately detect illegality and take appropriate countermeasures.

[0565] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0566] In this invention, the server includes means for a user to transmit video data to an information processing device via an information processing terminal, means for the information processing device to store the video data in a storage device and prepare it for analysis, and means for the information processing device to process the video data through a machine learning analysis device to perform video recognition and audio analysis. This makes it possible to evaluate the possibility of legal violations of the video data before posting it, and to make suggestions or control its publication as needed.

[0567] A "user" is the entity that transmits video data to the system via an information processing terminal.

[0568] An "information processing terminal" is a device used by a user to capture or edit video data and transmit it to an information processing device.

[0569] "Video data" refers to video content submitted by users, and its format and quality must be suitable for analysis on the server.

[0570] An "information processing device" is a computer system that receives transmitted video data and performs analysis and evaluation procedures.

[0571] A "storage device" is a digital storage device used by an information processing device to hold video data received by the device.

[0572] A "machine learning analysis device" is a technological device that performs image recognition and audio analysis on video data to evaluate the possibility of legal violations.

[0573] "Video recognition" is an analytical technology that identifies specific objects or people in a video and detects whether it contains illegal activities or infringing materials.

[0574] "Audio analysis" is a technology that analyzes the audio track within a video to determine whether certain music or sound effects infringe on any rights.

[0575] An "indicator for evaluating the possibility of legal violations" is an evaluation standard that quantifies the extent to which video data violates existing laws and regulations.

[0576] A "corporate entity" refers to an organization that operates the system and issues instructions regarding the publication of video data.

[0577] A "prompt message" is a set of instructions sent through a generative AI model to provide users or companies with legally-based improvement suggestions.

[0578] One embodiment of this invention is a system in which a user transmits video data to an information processing device using an information processing terminal, and based on this data, illegality is detected and managed. Specifically, it is implemented as follows.

[0579] First, the user captures or edits video data using their own information processing terminal. The terminal has video editing software installed, which allows the data to be converted and saved in an appropriate format. The prepared video data is then transmitted to the information processing device via the terminal.

[0580] The information processing device first saves the video data to a storage device. This storage device could be a large-capacity storage system or a cloud-based storage service. Afterward, this data is prepared to be passed to a machine learning analysis device.

[0581] The machine learning analysis system analyzes video data using video recognition and audio analysis technologies. This utilizes existing video analysis software and audio processing libraries. Video recognition identifies specific objects and actions, while audio analysis checks for copyright infringement of music and audio.

[0582] Based on the analysis results, the information processing device generates an index that quantifies the likelihood of legal violations. This index is analyzed by a generating AI model, and an anomaly is reported if it exceeds a certain threshold. The server notifies the company and also sends users legally-based improvement suggestions in the form of prompt messages.

[0583] For example, if copyrighted music is used in a video recorded by a user, the machine learning analysis system will detect this, and the server will suggest to the user that they delete or modify the audio portion. An example of a prompt used in this case is, "Please check if this video is illegal and suggest any parts that need to be corrected."

[0584] This system allows for the effective control of illegal content on the platform by checking the legality of video data before it is released and making timely improvements.

[0585] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0586] Step 1:

[0587] The user captures or edits video data using an information processing terminal.

[0588] Specific operation: The user uses dedicated video editing software to edit the recorded footage with filters and effects.

[0589] Input: Pre-recorded raw video data.

[0590] Data processing: Filtering and applying effects to the video using editing software.

[0591] Output: Video data file ready for submission.

[0592] Step 2:

[0593] The user transmits video data to the information processing device via an information processing terminal.

[0594] Specific operation: The user uses the upload function of the online platform to send video data to the server. Data format conversion is performed as needed.

[0595] Input: Edited video data file.

[0596] Data processing: Data compression and format conversion.

[0597] Output: Video data stored on the server.

[0598] Step 3:

[0599] The server saves the video data to a storage device and prepares it for analysis.

[0600] Specific operation: The server first saves the video data to secure storage and records metadata. It also performs the necessary preprocessing for analysis.

[0601] Input: Video data stored on the server.

[0602] Data processing: Extraction and recording of metadata.

[0603] Output: Stored video data and metadata.

[0604] Step 4:

[0605] The server uses a machine learning analysis system to perform image recognition and audio analysis.

[0606] Specific operation: The server uses video recognition algorithms to analyze objects and actions within the video, and audio analysis techniques to identify and analyze sound sources.

[0607] Input: Video data acquired from the storage device.

[0608] Data processing: Analysis of video frames and feature extraction from audio tracks.

[0609] Output: Analysis results of video and audio.

[0610] Step 5:

[0611] The server generates an index that evaluates the likelihood of legal violations based on the analysis results.

[0612] Specific operation: Based on the analysis results regarding specific legal violations, the server calculates evaluation metrics using a generated AI model and compares them to a threshold.

[0613] Input: Analysis results of video and audio.

[0614] Data processing: Integration of analysis results and calculation of evaluation metrics.

[0615] Output: Evaluation score for legal violations.

[0616] Step 6:

[0617] The server provides improvement suggestions and notifications to users and companies.

[0618] Specific actions: The server generates prompt messages and sends improvement suggestions based on legal grounds to the user. It also notifies the company of illegality.

[0619] Input: Evaluation metric scores and analysis results.

[0620] Data processing: Generating prompt messages and formatting notifications.

[0621] Output: Notifications and suggestions to users and organizations.

[0622] Step 7:

[0623] After user improvements are made, the server will re-analyze the video data and perform appropriate control.

[0624] Specific actions: The improved video data will be analyzed again to check if the problem has been resolved. If there are no problems, publication will be permitted.

[0625] Input: User-modified video data.

[0626] Data processing: Reanalysis and recalculation of evaluation metrics.

[0627] Output: Final evaluation results and decision on whether or not to publish.

[0628] (Application Example 1)

[0629] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0630] In content distribution services, there is a need for a system that can efficiently and reliably detect and address the legality of user-uploaded videos and images before distribution, especially if they may violate legal standards or regulations. However, existing systems primarily focus on post-distribution processing, leaving a risk that distributed content may be published while still containing illegal elements. Therefore, technology is needed to scrutinize the legality of content in advance and automatically take necessary measures.

[0631] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0632] In this invention, the server includes means for a user to transmit a video to a data processing device via an information processing device; means for the data processing device to hold the video in a data storage device and prepare it for analysis; means for the data processing device to run the video through a knowledge processing engine and perform visual recognition and audio analysis; and means for providing recommended alternative materials when a user modifies a video. This enables the legality of a video to be distributed to be detected in advance, and enables the distribution of safe and compliant content.

[0633] A "user" is an entity that transmits moving images using an information processing device and undergoes a process of verifying their legality.

[0634] An "information processing device" refers to a device used by a user to transmit video and image data to a data processing device.

[0635] "Motion images" refer to a series of images and audio data that include visual and auditory elements.

[0636] A "data processing unit" is a central system that receives transmitted video and performs various processes for analysis and storage.

[0637] A "data storage device" is a storage medium used to temporarily or permanently store moving images.

[0638] A "knowledge processing engine" is an algorithmic system that automatically analyzes the visual and auditory elements of moving images to determine their legality.

[0639] "Visual recognition" refers to the process of analyzing individual frames within a video or image to identify elements and patterns within that image.

[0640] "Audio analysis" is the process of analyzing audio data contained in video and determining its content and characteristics.

[0641] "Lawfulness" refers to the state in which a video or image does not violate any current laws or regulations.

[0642] A "numerical value" is a numerical evaluation indicating legality, and is an indicator of how problematic the video is in light of legal standards.

[0643] The "operating organization" is the organization responsible for managing and operating the content distribution service.

[0644] "Alternative material" refers to visual or auditory elements that can be used in place of the original video footage, as proposed to ensure legality.

[0645] The system for realizing this invention operates using a program with the following configuration.

[0646] The server receives video footage transmitted by the user using an information processing device. This video footage is temporarily stored in a data storage device. The server then transfers the stored video footage to a knowledge processing engine for visual recognition and audio analysis. This involves using multiple APIs that combine image and audio analysis technologies, with specific examples including Amazon Rekognition and Google Cloud Video Intelligence, depending on the use case.

[0647] Visual recognition analyzes video frames to identify image elements. Audio analysis analyzes audio tracks to identify their content. The server evaluates these analysis results and generates a numerical value indicating legality. This value is compared to current laws and regulations to determine how legal the video is.

[0648] If a user uploads a video containing illegal elements, the server will send the user a legally-based suggestion for improvement. This suggestion will include providing alternative material and showing how to correct or replace the problematic parts.

[0649] For example, if a video shot by a user on their smartphone contains copyrighted music, the server will recognize the music and suggest alternative music that can be used. A prompt message such as "This video contains copyrighted music. Please suggest alternative music that can be used" will be displayed to the user through the system, allowing the user to resolve legal issues and safely publish the video.

[0650] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0651] Step 1:

[0652] The user prepares video footage using an information processing device. The input here is video footage captured or edited by the user. To send this data to the data processing device, the video footage is encoded into an optimal format and prepared for transmission.

[0653] Step 2:

[0654] The terminal transmits video data to the data processing unit. The input is encoded video data, and the output is temporary storage in a data storage device by the data processing unit. Data is transmitted via a communication line, received by the data processing unit, and stored in the data storage device.

[0655] Step 3:

[0656] The server passes the video footage stored in the data storage device to the knowledge processing engine. The input is the stored video footage, and the output is data labeling in preparation for analysis. At this stage, the frames of the video footage are identified and each is converted into an analyzable format.

[0657] Step 4:

[0658] The server uses a knowledge processing engine to perform visual recognition and audio analysis of video. Inputs are labeled frames and audio data, and output is the analysis results based on each element. Visual recognition techniques identify objects and scenes within frames, while audio analysis techniques convert audio tracks into text and features. APIs such as Amazon Rekognition and Google Cloud Video Intelligence are used.

[0659] Step 5:

[0660] The server evaluates the legality of the video based on the analysis results. The input is the analysis results from step 4, and the output is a numerical score indicating the legality of the video. The score is compared to a pre-set legal standard to identify the frames and audio elements in which violations were detected.

[0661] Step 6:

[0662] The server sends the user legally-based improvement suggestions. The input is a legality score and a list of elements that need improvement, and the output is a notification and improvement suggestions for the user. Specifically, a prompt message is generated stating, "This video contains copyrighted music. Please recommend usable alternative music."

[0663] Step 7:

[0664] The user resends the improved video image to the server. The input is the corrected video image data, and the output is a notification that it is ready for re-analysis. The user corrects the video image based on the server's suggestions and resends it to complete the compliance verification process.

[0665] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0666] In this embodiment of the invention, a system is provided that, in addition to pre-evaluating video data, also has a function to analyze the user's emotions. The user uses a terminal to shoot or select a video and sends the video data to the server. Here, the server uses an emotion engine in addition to conventional analysis to analyze the user's emotions.

[0667] The server saves the video data to storage and prepares it for transfer to the AI ​​analysis engine. First, it performs image recognition and audio analysis to check for standard legal violations. Next, it uses an emotion engine to analyze how the video content and audio reflect the user's emotions. For example, it determines whether the user is expressing emotions such as anger, sadness, or joy based on their facial expressions and tone of voice in the video.

[0668] Once the analysis is complete, the server generates an illegality score and performs an overall evaluation that also takes into account additional sentiment analysis results. The sentiment analysis results contribute to determining whether the video content is intentionally illegal and provide additional information for notifications to the operating company and users.

[0669] The server then notifies the operating company of the illegality and sentiment analysis, and instructs them to take down or delete the video as necessary. The server also sends improvement suggestions to the user based on the analysis results, along with legal grounds and emotional feedback. This helps users not only correct illegality but also understand how their emotions are influencing the content. For example, if a user films a video while excited and it contains highly provocative content, the user will receive a message suggesting calm and constructive ways to improve the video.

[0670] The following describes the processing flow.

[0671] Step 1:

[0672] The user uses their device to record or select a video and sends the video data to the server. The user then enters basic information for posting and presses the submit button.

[0673] Step 2:

[0674] The server receives the video data and saves it to storage. During the saving process, the video metadata is extracted to prepare it for analysis.

[0675] Step 3:

[0676] The server activates the AI ​​analysis engine and first performs image recognition processing. This identifies potentially illegal elements from each frame in the video.

[0677] Step 4:

[0678] The server performs audio analysis and converts the audio in the video into text. This allows for the detection of potentially copyrighted music and illegal statements.

[0679] Step 5:

[0680] The server uses an emotion engine to analyze the user's emotions expressed in the video. It analyzes facial expressions, tone of voice, and other factors to determine the user's emotional state.

[0681] Step 6:

[0682] The server integrates the results of image recognition, audio analysis, and sentiment analysis to generate a video illegality score. In addition, it adjusts the overall score to reflect the sentiment analysis results.

[0683] Step 7:

[0684] The server will send a detailed notification to the operating company based on the illegality score and sentiment analysis results. Based on the information included, the operating company will consider whether or not to publish the video.

[0685] Step 8:

[0686] The server sends a notification of the analysis results to the user and provides suggestions for improvement. These suggestions include identifying illegal activities and providing feedback on sentiment, prompting the user to make appropriate corrections.

[0687] Step 9:

[0688] The server receives instructions from the operating company and, as necessary, takes the video data down, deletes it, or prepares it for re-uploading after correction. Videos corrected by the user are re-analyzed, and if the problem is resolved, re-uploading is permitted.

[0689] (Example 2)

[0690] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0691] In modern society, the rapid spread of video content has increased the risk of illegal or emotionally inappropriate content being published. This puts a strain on the resources of regulatory organizations and makes it difficult for users to make informed judgments about the legality and emotional impact of the content they are viewing.

[0692] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0693] In this invention, the server includes means for a user to transmit video information to a computer via a terminal, means for the computer to store the video information in a storage device and prepare it for analysis, and means for the computer to process the video information through an artificial intelligence analysis device to perform image recognition and audio analysis. This makes it possible to comprehensively analyze the illegality and emotional impact of video content, reduce the burden on management organizations, and provide useful feedback to users.

[0694] A "user" is an individual or group that manipulates video information via a terminal and transmits it to a computer.

[0695] A "terminal" is a device used by a user to capture or select video information.

[0696] "Video information" refers to video data that a user sends from their device to a computer.

[0697] A "computer" is a digital device that processes received video information and prepares it for analysis.

[0698] A "memory device" is a digital storage medium used to hold image information within a computer.

[0699] An "artificial intelligence analysis device" is a system that uses video information to perform image recognition and audio analysis.

[0700] "Image recognition" is the process of recognizing visual elements within video information and analyzing the data related to them.

[0701] "Audio analysis" is the process of analyzing the audio within video information and evaluating its content.

[0702] An "evaluation value" is a numerical value or indicator used to assess illegality based on the analysis results.

[0703] A "management organization" is an organization responsible for the appropriate handling and disclosure of video information.

[0704] A "notification" is information or an alert sent from a computer to an administrative organization or user.

[0705] "Improvement suggestions" are guidelines provided to users regarding specific changes or modifications to the content of video information.

[0706] "Feedback" refers to evaluations and opinions provided to users based on analysis results.

[0707] "Initial filtering" is the first process of selecting video information to be analyzed based on thresholds.

[0708] "Emotional analysis" is the process of analyzing a user's emotional state using visual and auditory elements within video information.

[0709] A "threshold" is a numerical value used as a criterion for selecting the target for analysis during the initial filtering process.

[0710] "Publication control" is the process of managing whether or not video information can be provided to viewers.

[0711] "Re-analysis" is the process of re-analyzing video information after the user has made improvements and evaluating the results.

[0712] This system begins with the user using a terminal to capture or select video information and sending that data to a server. The terminal includes camera functions and a file browser to assist with capturing and selecting video information.

[0713] The server stores the received video information in a storage device and prepares it for transfer to the artificial intelligence analysis device. Here, the storage device uses cloud storage or local server storage to securely hold the digital data. The server processes the video information through the artificial intelligence analysis device, which sequentially performs image recognition and audio analysis. Computer vision technology is used for image recognition to recognize objects and backgrounds in the video, while natural language processing technology is used for audio analysis to convert audio data into text and evaluate its content.

[0714] Furthermore, based on the analysis results, the server uses an emotion analysis engine to evaluate the user's emotional state. The emotion analysis engine incorporates facial recognition technology and voice tone analysis capabilities to analyze emotional elements within the video information in detail.

[0715] Once the analysis is complete, the server generates an evaluation score to assess the illegality and notifies the management organization of the results. At the same time, it sends users legally-based improvement suggestions and feedback that takes into account emotional impact, prompting them to take necessary action.

[0716] For example, if a user-created video content is deemed excessively offensive, the server will provide a suggestion for improvement such as, "This video contains offensive content. Please revise it to use milder language." Another example of a prompt for the generative AI model might be, "Analyze this video to determine its illegality or emotional impact, and generate appropriate feedback."

[0717] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0718] Step 1:

[0719] The user captures or selects video information using the device. The device acquires video information through its camera function and provides the ability to select the video file chosen by the user from local storage. It handles the video information selected by the user as input and generates data in a format to be sent to the server as output.

[0720] Step 2:

[0721] The terminal uploads the video information selected by the user to the server. During this process, it verifies that the video information is in the correct format and performs format conversion if necessary. The input is raw data stored on the terminal, and the output is digital data formatted for the server to receive.

[0722] Step 3:

[0723] The server stores the received video information in a storage device and prepares it for analysis. The server saves the files to secure cloud storage or local storage. The input is digital data transmitted from the terminal, and the output is a video file securely stored in the storage device.

[0724] Step 4:

[0725] The server inputs video information stored in its memory into an artificial intelligence analysis device, which then performs image recognition and audio analysis. Image recognition extracts visual elements, while audio analysis converts audio data into text and analyzes its content. The input is digital data including video and audio, and the output consists of identified image objects and analyzed audio text.

[0726] Step 5:

[0727] The server evaluates the user's emotional state using an emotion analysis engine based on the analyzed results. This involves using facial recognition technology to infer emotions from facial expressions in the video and analyzing the emotional elements of the voice using voice tone analysis. The input is the result data from image recognition and voice analysis, and the output is evaluation data regarding the user's emotional state.

[0728] Step 6:

[0729] The server generates an illegality rating based on the analysis results and sentiment evaluation, and sends a notification to the management organization. The rating is used as an indicator of the appropriateness of the video's publication. The input is all the analysis data, and the output is the notification to the management organization.

[0730] Step 7:

[0731] The server sends users legally-based improvement suggestions and provides feedback that takes emotional impact into consideration. The specific suggestions are customized based on the analysis results. Inputs are evaluation values ​​and analysis data, and output is a feedback message to the user.

[0732] (Application Example 2)

[0733] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0734] In modern information and communication technology, user-generated video content often contains inappropriate material, including legal violations, or may be deemed inappropriate due to the influence of user emotions. In such cases, there is a need for a system that provides appropriate filtering before video publication and constructive feedback to users. However, existing technologies lack sufficient means to analyze user emotions and conduct a comprehensive evaluation in conjunction with legal violations. Therefore, there is a need to provide a system that analyzes the legal and emotional aspects of video data and proactively controls and improves the transmission of inappropriate content.

[0735] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0736] In this invention, the server includes means for a user to transmit video data to an information processing device via an information processing device, means for the information processing device to analyze the user's facial expressions and voice and determine their emotions, and means for the information processing device to evaluate the influence of the user's emotional state on the content of the video data. This makes it possible to detect in advance if the video data is affected by legal violations or inappropriate emotions, and to provide appropriate feedback.

[0737] An "information processing device" is a computer system used for receiving, storing, analyzing, and notifying the results of data.

[0738] A "memory device" is a medium used by an information processing device to temporarily or permanently store video data.

[0739] A "data analysis device" is a system used by an information processing device to analyze video data, specifically for analyzing images and audio.

[0740] An "evaluation indicator" is a numerical value or indicator used for evaluation, generated based on legal compliance or other criteria.

[0741] An "organization" is an organization or group that receives notifications from information processing devices and issues instructions regarding the publication or processing of video data.

[0742] "User" refers to an individual or entity that transmits video data to an information processing device and receives analysis and feedback.

[0743] "Emotional discrimination" is the process of analyzing a user's facial expressions and voice to identify their emotional state.

[0744] "Initial filtering" refers to the process by which an information processing device performs basic screening before analyzing video data.

[0745] To implement this invention, a system centered on an information processing device is constructed. Its main components are as follows:

[0746] First, the user captures or selects video data using a smartphone or other device and sends it to the information processing device. This information processing device incorporates a storage device, a data analysis device, and the necessary software. Once the video data is received, the information processing device saves the data to its storage device and prepares it for analysis.

[0747] The data analysis device is equipped with an AI analysis engine for image recognition and speech analysis. Specifically, it uses the DeepFace library to analyze facial expressions and the SpeechRecognition library to analyze the content of speech data. Based on this data, the information processing device determines the user's emotional state and further evaluates how emotions influence the content of the video data.

[0748] Next, the information processing device generates evaluation indicators for legal violations and provides appropriate feedback to the user based on the evaluation results. This includes a function to judge the appropriateness of the video data based on the evaluation indicators and to provide specific improvement suggestions if improvements are needed.

[0749] For example, if a user films an event while excited, the system automatically determines whether the background audio and facial expressions are offensive. If offensive elements are detected, the information processing device sends feedback to the user, such as "We recommend editing this video to make it more positive," to encourage improvement.

[0750] Examples of prompts for generative AI models include the following:

[0751] "A user has requested an emotion analysis of the following video. Please analyze the facial expressions and tone of voice in the video to identify the main emotions. If the emotions are negative, please provide specific advice on how to correct them."

[0752] In this way, it becomes possible to provide comprehensive evaluation and improvement support for video data via information processing equipment.

[0753] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0754] Step 1:

[0755] The user uses a device to capture or select video data. The input is video data from the user's device. The device sends this data to the information processing device. The output is the video data sent to the server.

[0756] Step 2:

[0757] The server saves the received video data to its storage device. The input is video data transmitted from the terminal, which the server temporarily stores in its storage device. The output is data ready for analysis.

[0758] Step 3:

[0759] The server passes the video data stored in its memory to the AI ​​analysis engine, initiating image recognition and audio analysis. The input is the stored video data. The server then uses the DeepFace library to analyze facial expressions and the SpeechRecognition library to convert audio data into text. The output consists of the user's facial expression data and the content of the audio.

[0760] Step 4:

[0761] The server determines the user's emotions based on the analysis results. The input consists of image recognition and voice analysis results, which the server uses to estimate the user's emotions. The output provides the user's main emotions (e.g., joy, anger, sadness).

[0762] Step 5:

[0763] The server evaluates the legal compliance and emotional state of the video data and generates an evaluation index. Inputs include a checklist of legal violations and the results of the emotional assessment. Based on this, the server generates an evaluation index that quantifies the safety and appropriateness of the video data. The output is the evaluation index.

[0764] Step 6:

[0765] The server provides feedback to the user based on evaluation metrics it generates. Inputs include evaluation metrics and improvement suggestion templates. The server generates prompt messages, creates specific feedback such as "We recommend editing this video to be more positive," and sends it to the user. Output is the feedback message sent to the user.

[0766] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0767] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0768] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0769] [Fourth Embodiment]

[0770] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0771] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0772] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0773] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0774] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0775] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0776] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0777] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0778] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0779] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0780] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0781] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0782] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0783] One embodiment of this invention is a system in which a user uses a terminal to send video data to a server, and a series of illegality detection processes are provided based on that data.

[0784] First, the user shoots or edits video data on their device and prepares to post it to the platform. When the user posts, the video data is sent to the server.

[0785] The server stores the received video data in storage and prepares it for the AI ​​analysis engine. The video data is analyzed using image recognition technology to identify elements of still images, and audio analysis technology to identify the content and characteristics of the sound.

[0786] Next, the server uses an AI analysis engine to analyze each frame in the video to check for any elements that violate specific laws or regulations. For example, it determines whether copyrighted music is being used or whether illegal acts or items are depicted.

[0787] Based on the analysis, the server generates a score that evaluates the illegality of the video data. This score indicates the extent to which the video violates existing laws, and if it exceeds a certain threshold, it is deemed illegal.

[0788] Subsequently, the server sends a notification to the operating company regarding the video deemed illegal. The operating company receives this notification and sends instructions to the server regarding whether or not to stop or delete the video.

[0789] Furthermore, the server will also notify users about the illegality and offer suggestions for improvement. These suggestions may include methods for deleting or modifying specific parts of the video.

[0790] Furthermore, based on instructions from the operating company, the server controls the publication of problematic video data, re-analyzes it after user corrections, and allows video posting if appropriate.

[0791] This system allows for the detection and management of illegal video content before it is uploaded, effectively suppressing illegal content on the platform. For example, if illegal use of music is detected in a video uploaded by a user, the system recommends that the offending music be removed, and the upload is permitted only after the removal process is completed.

[0792] The following describes the processing flow.

[0793] Step 1:

[0794] The user uses their device to record or select a video and sends the video data to the server for posting. The video data also includes metadata.

[0795] Step 2:

[0796] The server receives the video data and saves it to storage. This saving is done in preparation for analysis, and data integrity is checked simultaneously.

[0797] Step 3:

[0798] The server passes the video to the AI ​​analysis engine, which then begins image recognition processing. Specifically, it detects specific objects and text from each frame in the video and checks for items that may be legally problematic.

[0799] Step 4:

[0800] The server continues to perform audio analysis using its AI analysis engine. It converts the audio contained in the video into text and detects the use of copyrighted music and illegal words.

[0801] Step 5:

[0802] The server integrates the results of image recognition and audio analysis to generate a score that assesses illegality. This score is calculated based on the importance of each element that violates the law.

[0803] Step 6:

[0804] Based on the illegality score generated by the server, if a certain threshold is exceeded, the operating company will be notified. The notification will include detailed analysis results.

[0805] Step 7:

[0806] The server sends the user a notification based on the analysis results, offering suggestions for improving problematic content or audio. The user then considers correcting the video based on these suggestions.

[0807] Step 8:

[0808] Based on instructions from the operating company, the server will decide to take down, delete, or partially edit the content of the video in question and perform the necessary actions.

[0809] Step 9:

[0810] After the user corrects the video based on the server's notification, the server re-analyzes it. Once it confirms that the problem has been resolved, the server allows the video to be posted or made public.

[0811] (Example 1)

[0812] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0813] Currently, the publication of illegal content in video data posted on online platforms is a serious problem. In particular, videos containing copyright infringement or inappropriate content are often published without prior detection, increasing legal risks for platform users and operators. To solve this problem, it is necessary to efficiently and accurately detect illegality and take appropriate countermeasures.

[0814] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0815] In this invention, the server includes means for a user to transmit video data to an information processing device via an information processing terminal, means for the information processing device to store the video data in a storage device and prepare it for analysis, and means for the information processing device to process the video data through a machine learning analysis device to perform video recognition and audio analysis. This makes it possible to evaluate the possibility of legal violations of the video data before posting it, and to make suggestions or control its publication as needed.

[0816] A "user" is the entity that transmits video data to the system via an information processing terminal.

[0817] An "information processing terminal" is a device used by a user to capture or edit video data and transmit it to an information processing device.

[0818] "Video data" refers to video content submitted by users, and its format and quality must be suitable for analysis on the server.

[0819] An "information processing device" is a computer system that receives transmitted video data and performs analysis and evaluation procedures.

[0820] A "storage device" is a digital storage device used by an information processing device to hold video data received by the device.

[0821] A "machine learning analysis device" is a technological device that performs image recognition and audio analysis on video data to evaluate the possibility of legal violations.

[0822] "Video recognition" is an analytical technology that identifies specific objects or people in a video and detects whether it contains illegal activities or infringing materials.

[0823] "Audio analysis" is a technology that analyzes the audio track within a video to determine whether certain music or sound effects infringe on any rights.

[0824] An "indicator for evaluating the possibility of legal violations" is an evaluation standard that quantifies the extent to which video data violates existing laws and regulations.

[0825] A "corporate entity" refers to an organization that operates the system and issues instructions regarding the publication of video data.

[0826] A "prompt message" is a set of instructions sent through a generative AI model to provide users or companies with legally-based improvement suggestions.

[0827] One embodiment of this invention is a system in which a user transmits video data to an information processing device using an information processing terminal, and based on this data, illegality is detected and managed. Specifically, it is implemented as follows.

[0828] First, the user captures or edits video data using their own information processing terminal. The terminal has video editing software installed, which allows the data to be converted and saved in an appropriate format. The prepared video data is then transmitted to the information processing device via the terminal.

[0829] The information processing device first saves the video data to a storage device. This storage device could be a large-capacity storage system or a cloud-based storage service. Afterward, this data is prepared to be passed to a machine learning analysis device.

[0830] The machine learning analysis system analyzes video data using video recognition and audio analysis technologies. This utilizes existing video analysis software and audio processing libraries. Video recognition identifies specific objects and actions, while audio analysis checks for copyright infringement of music and audio.

[0831] Based on the analysis results, the information processing device generates an index that quantifies the likelihood of legal violations. This index is analyzed by a generating AI model, and an anomaly is reported if it exceeds a certain threshold. The server notifies the company and also sends users legally-based improvement suggestions in the form of prompt messages.

[0832] For example, if copyrighted music is used in a video recorded by a user, the machine learning analysis system will detect this, and the server will suggest to the user that they delete or modify the audio portion. An example of a prompt used in this case is, "Please check if this video is illegal and suggest any parts that need to be corrected."

[0833] This system allows for the effective control of illegal content on the platform by checking the legality of video data before it is released and making timely improvements.

[0834] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0835] Step 1:

[0836] The user captures or edits video data using an information processing terminal.

[0837] Specific operation: The user uses dedicated video editing software to edit the recorded footage with filters and effects.

[0838] Input: Pre-recorded raw video data.

[0839] Data processing: Filtering and applying effects to the video using editing software.

[0840] Output: Video data file ready for submission.

[0841] Step 2:

[0842] The user transmits video data to the information processing device via an information processing terminal.

[0843] Specific operation: The user uses the upload function of the online platform to send video data to the server. Data format conversion is performed as needed.

[0844] Input: Edited video data file.

[0845] Data processing: Data compression and format conversion.

[0846] Output: Video data stored on the server.

[0847] Step 3:

[0848] The server saves the video data to a storage device and prepares it for analysis.

[0849] Specific operation: The server first saves the video data to secure storage and records metadata. It also performs the necessary preprocessing for analysis.

[0850] Input: Video data stored on the server.

[0851] Data processing: Extraction and recording of metadata.

[0852] Output: Stored video data and metadata.

[0853] Step 4:

[0854] The server uses a machine learning analysis system to perform image recognition and audio analysis.

[0855] Specific operation: The server uses video recognition algorithms to analyze objects and actions within the video, and audio analysis techniques to identify and analyze sound sources.

[0856] Input: Video data acquired from the storage device.

[0857] Data processing: Analysis of video frames and feature extraction from audio tracks.

[0858] Output: Analysis results of video and audio.

[0859] Step 5:

[0860] The server generates an index that evaluates the likelihood of legal violations based on the analysis results.

[0861] Specific operation: Based on the analysis results regarding specific legal violations, the server calculates evaluation metrics using a generated AI model and compares them to a threshold.

[0862] Input: Analysis results of video and audio.

[0863] Data processing: Integration of analysis results and calculation of evaluation metrics.

[0864] Output: Evaluation score for legal violations.

[0865] Step 6:

[0866] The server provides improvement suggestions and notifications to users and companies.

[0867] Specific actions: The server generates prompt messages and sends improvement suggestions based on legal grounds to the user. It also notifies the company of illegality.

[0868] Input: Evaluation metric scores and analysis results.

[0869] Data processing: Generating prompt messages and formatting notifications.

[0870] Output: Notifications and suggestions to users and organizations.

[0871] Step 7:

[0872] After user improvements are made, the server will re-analyze the video data and perform appropriate control.

[0873] Specific actions: The improved video data will be analyzed again to check if the problem has been resolved. If there are no problems, publication will be permitted.

[0874] Input: User-modified video data.

[0875] Data processing: Reanalysis and recalculation of evaluation metrics.

[0876] Output: Final evaluation results and decision on whether or not to publish.

[0877] (Application Example 1)

[0878] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0879] In content distribution services, there is a need for a system that can efficiently and reliably detect and address the legality of user-uploaded videos and images before distribution, especially if they may violate legal standards or regulations. However, existing systems primarily focus on post-distribution processing, leaving a risk that distributed content may be published while still containing illegal elements. Therefore, technology is needed to scrutinize the legality of content in advance and automatically take necessary measures.

[0880] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0881] In this invention, the server includes means for a user to transmit a video to a data processing device via an information processing device; means for the data processing device to hold the video in a data storage device and prepare it for analysis; means for the data processing device to run the video through a knowledge processing engine and perform visual recognition and audio analysis; and means for providing recommended alternative materials when a user modifies a video. This enables the legality of a video to be distributed to be detected in advance, and enables the distribution of safe and compliant content.

[0882] A "user" is an entity that transmits moving images using an information processing device and undergoes a process of verifying their legality.

[0883] An "information processing device" refers to a device used by a user to transmit video and image data to a data processing device.

[0884] "Motion images" refer to a series of images and audio data that include visual and auditory elements.

[0885] A "data processing unit" is a central system that receives transmitted video and performs various processes for analysis and storage.

[0886] A "data storage device" is a storage medium used to temporarily or permanently store moving images.

[0887] A "knowledge processing engine" is an algorithmic system that automatically analyzes the visual and auditory elements of moving images to determine their legality.

[0888] "Visual recognition" refers to the process of analyzing individual frames within a video or image to identify elements and patterns within that image.

[0889] "Audio analysis" is the process of analyzing audio data contained in video and determining its content and characteristics.

[0890] "Lawfulness" refers to the state in which a video or image does not violate any current laws or regulations.

[0891] A "numerical value" is a numerical evaluation indicating legality, and is an indicator of how problematic the video is in light of legal standards.

[0892] The "operating organization" is the organization responsible for managing and operating the content distribution service.

[0893] "Alternative material" refers to visual or auditory elements that can be used in place of the original video footage, as proposed to ensure legality.

[0894] The system for realizing this invention operates using a program with the following configuration.

[0895] The server receives video footage transmitted by the user using an information processing device. This video footage is temporarily stored in a data storage device. The server then transfers the stored video footage to a knowledge processing engine for visual recognition and audio analysis. This involves using multiple APIs that combine image and audio analysis technologies, with specific examples including Amazon Rekognition and Google Cloud Video Intelligence, depending on the use case.

[0896] Visual recognition analyzes video frames to identify image elements. Audio analysis analyzes audio tracks to identify their content. The server evaluates these analysis results and generates a numerical value indicating legality. This value is compared to current laws and regulations to determine how legal the video is.

[0897] If a user uploads a video containing illegal elements, the server will send the user a legally-based suggestion for improvement. This suggestion will include providing alternative material and showing how to correct or replace the problematic parts.

[0898] For example, if a video shot by a user on their smartphone contains copyrighted music, the server will recognize the music and suggest alternative music that can be used. A prompt message such as "This video contains copyrighted music. Please suggest alternative music that can be used" will be displayed to the user through the system, allowing the user to resolve legal issues and safely publish the video.

[0899] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0900] Step 1:

[0901] The user prepares video footage using an information processing device. The input here is video footage captured or edited by the user. To send this data to the data processing device, the video footage is encoded into an optimal format and prepared for transmission.

[0902] Step 2:

[0903] The terminal transmits video data to the data processing unit. The input is encoded video data, and the output is temporary storage in a data storage device by the data processing unit. Data is transmitted via a communication line, received by the data processing unit, and stored in the data storage device.

[0904] Step 3:

[0905] The server passes the video footage stored in the data storage device to the knowledge processing engine. The input is the stored video footage, and the output is data labeling in preparation for analysis. At this stage, the frames of the video footage are identified and each is converted into an analyzable format.

[0906] Step 4:

[0907] The server uses a knowledge processing engine to perform visual recognition and audio analysis of video. Inputs are labeled frames and audio data, and output is the analysis results based on each element. Visual recognition techniques identify objects and scenes within frames, while audio analysis techniques convert audio tracks into text and features. APIs such as Amazon Rekognition and Google Cloud Video Intelligence are used.

[0908] Step 5:

[0909] The server evaluates the legality of the video based on the analysis results. The input is the analysis results from step 4, and the output is a numerical score indicating the legality of the video. The score is compared to a pre-set legal standard to identify the frames and audio elements in which violations were detected.

[0910] Step 6:

[0911] The server sends the user legally-based improvement suggestions. The input is a legality score and a list of elements that need improvement, and the output is a notification and improvement suggestions for the user. Specifically, a prompt message is generated stating, "This video contains copyrighted music. Please recommend usable alternative music."

[0912] Step 7:

[0913] The user resends the improved video image to the server. The input is the corrected video image data, and the output is a notification that it is ready for re-analysis. The user corrects the video image based on the server's suggestions and resends it to complete the compliance verification process.

[0914] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0915] In this embodiment of the invention, a system is provided that, in addition to pre-evaluating video data, also has a function to analyze the user's emotions. The user uses a terminal to shoot or select a video and sends the video data to the server. Here, the server uses an emotion engine in addition to conventional analysis to analyze the user's emotions.

[0916] The server saves the video data to storage and prepares it for transfer to the AI ​​analysis engine. First, it performs image recognition and audio analysis to check for standard legal violations. Next, it uses an emotion engine to analyze how the video content and audio reflect the user's emotions. For example, it determines whether the user is expressing emotions such as anger, sadness, or joy based on their facial expressions and tone of voice in the video.

[0917] Once the analysis is complete, the server generates an illegality score and performs an overall evaluation that also takes into account additional sentiment analysis results. The sentiment analysis results contribute to determining whether the video content is intentionally illegal and provide additional information for notifications to the operating company and users.

[0918] The server then notifies the operating company of the illegality and sentiment analysis, and instructs them to take down or delete the video as necessary. The server also sends improvement suggestions to the user based on the analysis results, along with legal grounds and emotional feedback. This helps users not only correct illegality but also understand how their emotions are influencing the content. For example, if a user films a video while excited and it contains highly provocative content, the user will receive a message suggesting calm and constructive ways to improve the video.

[0919] The following describes the processing flow.

[0920] Step 1:

[0921] The user uses their device to record or select a video and sends the video data to the server. The user then enters basic information for posting and presses the submit button.

[0922] Step 2:

[0923] The server receives the video data and saves it to storage. During the saving process, the video metadata is extracted to prepare it for analysis.

[0924] Step 3:

[0925] The server activates the AI ​​analysis engine and first performs image recognition processing. This identifies potentially illegal elements from each frame in the video.

[0926] Step 4:

[0927] The server performs audio analysis and converts the audio in the video into text. This allows for the detection of potentially copyrighted music and illegal statements.

[0928] Step 5:

[0929] The server uses an emotion engine to analyze the user's emotions expressed in the video. It analyzes facial expressions, tone of voice, and other factors to determine the user's emotional state.

[0930] Step 6:

[0931] The server integrates the results of image recognition, audio analysis, and sentiment analysis to generate a video illegality score. In addition, it adjusts the overall score to reflect the sentiment analysis results.

[0932] Step 7:

[0933] The server will send a detailed notification to the operating company based on the illegality score and sentiment analysis results. Based on the information included, the operating company will consider whether or not to publish the video.

[0934] Step 8:

[0935] The server sends a notification of the analysis results to the user and provides suggestions for improvement. These suggestions include identifying illegal activities and providing feedback on sentiment, prompting the user to make appropriate corrections.

[0936] Step 9:

[0937] The server receives instructions from the operating company and, as necessary, takes the video data down, deletes it, or prepares it for re-uploading after correction. Videos corrected by the user are re-analyzed, and if the problem is resolved, re-uploading is permitted.

[0938] (Example 2)

[0939] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0940] In modern society, the rapid spread of video content has increased the risk of illegal or emotionally inappropriate content being published. This puts a strain on the resources of regulatory organizations and makes it difficult for users to make informed judgments about the legality and emotional impact of the content they are viewing.

[0941] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0942] In this invention, the server includes means for a user to transmit video information to a computer via a terminal, means for the computer to store the video information in a storage device and prepare it for analysis, and means for the computer to process the video information through an artificial intelligence analysis device to perform image recognition and audio analysis. This makes it possible to comprehensively analyze the illegality and emotional impact of video content, reduce the burden on management organizations, and provide useful feedback to users.

[0943] A "user" is an individual or group that manipulates video information via a terminal and transmits it to a computer.

[0944] A "terminal" is a device used by a user to capture or select video information.

[0945] "Video information" refers to video data that a user sends from their device to a computer.

[0946] A "computer" is a digital device that processes received video information and prepares it for analysis.

[0947] A "memory device" is a digital storage medium used to hold image information within a computer.

[0948] An "artificial intelligence analysis device" is a system that uses video information to perform image recognition and audio analysis.

[0949] "Image recognition" is the process of recognizing visual elements within video information and analyzing the data related to them.

[0950] "Audio analysis" is the process of analyzing the audio within video information and evaluating its content.

[0951] An "evaluation value" is a numerical value or indicator used to assess illegality based on the analysis results.

[0952] A "management organization" is an organization responsible for the appropriate handling and disclosure of video information.

[0953] A "notification" is information or an alert sent from a computer to an administrative organization or user.

[0954] "Improvement suggestions" are guidelines provided to users regarding specific changes or modifications to the content of video information.

[0955] "Feedback" refers to evaluations and opinions provided to users based on analysis results.

[0956] "Initial filtering" is the first process of selecting video information to be analyzed based on thresholds.

[0957] "Emotional analysis" is the process of analyzing a user's emotional state using visual and auditory elements within video information.

[0958] A "threshold" is a numerical value used as a criterion for selecting the target for analysis during the initial filtering process.

[0959] "Publication control" is the process of managing whether or not video information can be provided to viewers.

[0960] "Re-analysis" is the process of re-analyzing video information after the user has made improvements and evaluating the results.

[0961] This system begins with the user using a terminal to capture or select video information and sending that data to a server. The terminal includes camera functions and a file browser to assist with capturing and selecting video information.

[0962] The server stores the received video information in a storage device and prepares it for transfer to the artificial intelligence analysis device. Here, the storage device uses cloud storage or local server storage to securely hold the digital data. The server processes the video information through the artificial intelligence analysis device, which sequentially performs image recognition and audio analysis. Computer vision technology is used for image recognition to recognize objects and backgrounds in the video, while natural language processing technology is used for audio analysis to convert audio data into text and evaluate its content.

[0963] Furthermore, based on the analysis results, the server uses an emotion analysis engine to evaluate the user's emotional state. The emotion analysis engine incorporates facial recognition technology and voice tone analysis capabilities to analyze emotional elements within the video information in detail.

[0964] Once the analysis is complete, the server generates an evaluation score to assess the illegality and notifies the management organization of the results. At the same time, it sends users legally-based improvement suggestions and feedback that takes into account emotional impact, prompting them to take necessary action.

[0965] For example, if a user-created video content is deemed excessively offensive, the server will provide a suggestion for improvement such as, "This video contains offensive content. Please revise it to use milder language." Another example of a prompt for the generative AI model might be, "Analyze this video to determine its illegality or emotional impact, and generate appropriate feedback."

[0966] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0967] Step 1:

[0968] The user captures or selects video information using the device. The device acquires video information through its camera function and provides the ability to select the video file chosen by the user from local storage. It handles the video information selected by the user as input and generates data in a format to be sent to the server as output.

[0969] Step 2:

[0970] The terminal uploads the video information selected by the user to the server. During this process, it verifies that the video information is in the correct format and performs format conversion if necessary. The input is raw data stored on the terminal, and the output is digital data formatted for the server to receive.

[0971] Step 3:

[0972] The server stores the received video information in a storage device and prepares it for analysis. The server saves the files to secure cloud storage or local storage. The input is digital data transmitted from the terminal, and the output is a video file securely stored in the storage device.

[0973] Step 4:

[0974] The server inputs video information stored in its memory into an artificial intelligence analysis device, which then performs image recognition and audio analysis. Image recognition extracts visual elements, while audio analysis converts audio data into text and analyzes its content. The input is digital data including video and audio, and the output consists of identified image objects and analyzed audio text.

[0975] Step 5:

[0976] The server evaluates the user's emotional state using an emotion analysis engine based on the analyzed results. This involves using facial recognition technology to infer emotions from facial expressions in the video and analyzing the emotional elements of the voice using voice tone analysis. The input is the result data from image recognition and voice analysis, and the output is evaluation data regarding the user's emotional state.

[0977] Step 6:

[0978] The server generates an illegality rating based on the analysis results and sentiment evaluation, and sends a notification to the management organization. The rating is used as an indicator of the appropriateness of the video's publication. The input is all the analysis data, and the output is the notification to the management organization.

[0979] Step 7:

[0980] The server sends users legally-based improvement suggestions and provides feedback that takes emotional impact into consideration. The specific suggestions are customized based on the analysis results. Inputs are evaluation values ​​and analysis data, and output is a feedback message to the user.

[0981] (Application Example 2)

[0982] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0983] In modern information and communication technology, user-generated video content often contains inappropriate material, including legal violations, or may be deemed inappropriate due to the influence of user emotions. In such cases, there is a need for a system that provides appropriate filtering before video publication and constructive feedback to users. However, existing technologies lack sufficient means to analyze user emotions and conduct a comprehensive evaluation in conjunction with legal violations. Therefore, there is a need to provide a system that analyzes the legal and emotional aspects of video data and proactively controls and improves the transmission of inappropriate content.

[0984] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0985] In this invention, the server includes means for a user to transmit video data to an information processing device via an information processing device, means for the information processing device to analyze the user's facial expressions and voice and determine their emotions, and means for the information processing device to evaluate the influence of the user's emotional state on the content of the video data. This makes it possible to detect in advance if the video data is affected by legal violations or inappropriate emotions, and to provide appropriate feedback.

[0986] An "information processing device" is a computer system used for receiving, storing, analyzing, and notifying the results of data.

[0987] A "memory device" is a medium used by an information processing device to temporarily or permanently store video data.

[0988] A "data analysis device" is a system used by an information processing device to analyze video data, specifically for analyzing images and audio.

[0989] An "evaluation indicator" is a numerical value or indicator used for evaluation, generated based on legal compliance or other criteria.

[0990] An "organization" is an organization or group that receives notifications from information processing devices and issues instructions regarding the publication or processing of video data.

[0991] "User" refers to an individual or entity that transmits video data to an information processing device and receives analysis and feedback.

[0992] "Emotional discrimination" is the process of analyzing a user's facial expressions and voice to identify their emotional state.

[0993] "Initial filtering" refers to the process by which an information processing device performs basic screening before analyzing video data.

[0994] To implement this invention, a system centered on an information processing device is constructed. Its main components are as follows:

[0995] First, the user captures or selects video data using a smartphone or other device and sends it to the information processing device. This information processing device incorporates a storage device, a data analysis device, and the necessary software. Once the video data is received, the information processing device saves the data to its storage device and prepares it for analysis.

[0996] The data analysis device is equipped with an AI analysis engine for image recognition and speech analysis. Specifically, it uses the DeepFace library to analyze facial expressions and the SpeechRecognition library to analyze the content of speech data. Based on this data, the information processing device determines the user's emotional state and further evaluates how emotions influence the content of the video data.

[0997] Next, the information processing device generates evaluation indicators for legal violations and provides appropriate feedback to the user based on the evaluation results. This includes a function to judge the appropriateness of the video data based on the evaluation indicators and to provide specific improvement suggestions if improvements are needed.

[0998] For example, if a user films an event while excited, the system automatically determines whether the background audio and facial expressions are offensive. If offensive elements are detected, the information processing device sends feedback to the user, such as "We recommend editing this video to make it more positive," to encourage improvement.

[0999] Examples of prompts for generative AI models include the following:

[1000] "A user has requested an emotion analysis of the following video. Please analyze the facial expressions and tone of voice in the video to identify the main emotions. If the emotions are negative, please provide specific advice on how to correct them."

[1001] In this way, it becomes possible to provide comprehensive evaluation and improvement support for video data via information processing equipment.

[1002] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1003] Step 1:

[1004] The user uses a device to capture or select video data. The input is video data from the user's device. The device sends this data to the information processing device. The output is the video data sent to the server.

[1005] Step 2:

[1006] The server saves the received video data to its storage device. The input is video data transmitted from the terminal, which the server temporarily stores in its storage device. The output is data ready for analysis.

[1007] Step 3:

[1008] The server passes the video data stored in its memory to the AI ​​analysis engine, initiating image recognition and audio analysis. The input is the stored video data. The server then uses the DeepFace library to analyze facial expressions and the SpeechRecognition library to convert audio data into text. The output consists of the user's facial expression data and the content of the audio.

[1009] Step 4:

[1010] The server determines the user's emotions based on the analysis results. The input consists of image recognition and voice analysis results, which the server uses to estimate the user's emotions. The output provides the user's main emotions (e.g., joy, anger, sadness).

[1011] Step 5:

[1012] The server evaluates the legal compliance and emotional state of the video data and generates an evaluation index. Inputs include a checklist of legal violations and the results of the emotional assessment. Based on this, the server generates an evaluation index that quantifies the safety and appropriateness of the video data. The output is the evaluation index.

[1013] Step 6:

[1014] The server provides feedback to the user based on evaluation metrics it generates. Inputs include evaluation metrics and improvement suggestion templates. The server generates prompt messages, creates specific feedback such as "We recommend editing this video to be more positive," and sends it to the user. Output is the feedback message sent to the user.

[1015] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1016] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1017] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1018] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1019] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1020] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1021] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1022] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1023] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1024] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1025] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1026] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1027] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1028] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1029] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1030] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1031] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1032] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1033] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1034] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1035] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1036] The following is further disclosed regarding the embodiments described above.

[1037] (Claim 1)

[1038] A means for a user to send video data to a server via their device,

[1039] The server has means for storing the video data in storage and preparing it for analysis,

[1040] A server processes video data using an AI analysis engine to perform image recognition and audio analysis.

[1041] A means for a server to generate a score that evaluates illegality based on the analysis results,

[1042] A means for the server to notify the operating company based on its illegality score,

[1043] A means by which the server sends improvement suggestions to the user based on legal grounds,

[1044] Based on instructions from the operating company, the server has a means to process video data appropriately,

[1045] A system that includes this.

[1046] (Claim 2)

[1047] The system according to claim 1, comprising a server that performs initial filtering of video data and means for selecting data to be analyzed based on a threshold.

[1048] (Claim 3)

[1049] The system according to claim 1, comprising means for appropriately controlling the publication of video data based on instructions received from the operating company, and for re-analyzing the data after the user has made improvements.

[1050] "Example 1"

[1051] (Claim 1)

[1052] A means by which a user transmits video data to an information processing device via an information processing terminal,

[1053] The information processing device holds the video data in a storage device and prepares it for analysis,

[1054] The information processing device processes video data through a machine learning analysis device, and the means for performing video recognition and audio analysis are provided.

[1055] A means by which an information processing device generates an index for evaluating the possibility of legal violations based on the analysis results,

[1056] A means by which an information processing device notifies a company based on indicators of illegality,

[1057] A means by which an information processing device sends improvement suggestions to a user based on legal grounds,

[1058] Based on instructions from a corporate entity, the information processing device provides a means for appropriately processing video data,

[1059] A means of re-evaluating problems in user-provided video data using a machine learning analysis device,

[1060] A system that includes this.

[1061] (Claim 2)

[1062] The system according to claim 1, wherein the information processing device has means for performing initial filtering of video data and selecting targets for analysis based on thresholds.

[1063] (Claim 3)

[1064] The system according to claim 1, comprising means for appropriately controlling the publication of video data based on instructions received from a corporate entity and for performing re-analysis after the user has made improvements.

[1065] "Application Example 1"

[1066] (Claim 1)

[1067] A means by which a user transmits a moving image to a data processing device via an information processing device,

[1068] The data processing device includes means for storing the video image in a data storage device and preparing it for analysis,

[1069] A data processing device processes moving images into a knowledge processing engine and performs visual recognition and audio analysis.

[1070] A data processing device provides means for generating numerical values ​​to evaluate legality based on analysis results,

[1071] A means by which a data processing device notifies the operating organization based on numerical values ​​of legality,

[1072] A means by which a data processing device sends improvement suggestions based on legal grounds to a user,

[1073] Based on instructions from the operating organization, the data processing unit has a means to appropriately process the video images,

[1074] When users modify video images, a means of providing recommended alternative materials,

[1075] A system that includes this.

[1076] (Claim 2)

[1077] The system according to claim 1, wherein the data processing device has means for performing initial filtering of moving images and selecting targets for analysis based on thresholds.

[1078] (Claim 3)

[1079] The system according to claim 1, comprising means for appropriately controlling the publication of video images based on instructions received from the operating organization, and for performing re-analysis after the user has made improvements.

[1080] "Example 2 of combining an emotion engine"

[1081] (Claim 1)

[1082] A means for a user to transmit video information to a computer via a terminal,

[1083] A means by which a computer stores the video information in a storage device and prepares it for analysis,

[1084] A means by which a computer processes video information through an artificial intelligence analysis device to perform image recognition and audio analysis,

[1085] A means by which a computer generates an evaluation value that assesses illegality based on the analysis results,

[1086] A means by which a computer notifies the management organization based on an assessment of illegality,

[1087] A means by which a computer sends users legally-based improvement suggestions and provides feedback that takes emotional impact into consideration,

[1088] Based on instructions from the management organization, the means by which the computer appropriately processes video information,

[1089] A system that includes this.

[1090] (Claim 2)

[1091] The system according to claim 1, comprising means for a computer to perform initial filtering of video information, select subjects for analysis based on thresholds, and perform emotion analysis to evaluate the user's emotional state.

[1092] (Claim 3)

[1093] The system according to claim 1, comprising means for appropriately controlling the disclosure of video information based on instructions received from a management organization, re-analyzing the data after the user has made improvements, and evaluating changes in emotional state.

[1094] "Application example 2 when combining with an emotional engine"

[1095] (Claim 1)

[1096] A means by which a user transmits video data to an information processing device via an information processing device,

[1097] The information processing device has means for storing the video data in a storage device and preparing it for analysis,

[1098] The information processing device processes video data through a data analysis device, and the means for performing image analysis and audio analysis are provided.

[1099] A means by which an information processing device generates an evaluation index for evaluating legal violations based on the analysis results,

[1100] A means by which an information processing device notifies an organization based on an evaluation index for legal violations,

[1101] A means by which an information processing device sends improvement suggestions based on legal grounds to a user,

[1102] Based on instructions from the organization, the information processing device has means to appropriately process video data,

[1103] The information processing device analyzes the user's facial expressions and voice to determine their emotions,

[1104] A means for an information processing device to evaluate the influence of a user's emotional state on the content of video data,

[1105] A system that includes this.

[1106] (Claim 2)

[1107] The system according to claim 1, wherein the information processing device has means for performing initial filtering of video data and selecting targets for analysis based on thresholds.

[1108] (Claim 3)

[1109] The system according to claim 1, comprising means for appropriately controlling the publication of video data based on instructions received from an organization, and for re-analyzing the data after improvements have been made by the user. [Explanation of symbols]

[1110] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for a user to send video data to a server via their device, The server has means for storing the video data in storage and preparing it for analysis, A server processes video data using an AI analysis engine to perform image recognition and audio analysis. A means for a server to generate a score that evaluates illegality based on the analysis results, A means for the server to notify the operating company based on its illegality score, A means by which the server sends improvement suggestions to the user based on legal grounds, Based on instructions from the operating company, the server has a means to process video data appropriately, A system that includes this.

2. The system according to claim 1, comprising a server that performs initial filtering of video data and means for selecting data to be analyzed based on a threshold.

3. The system according to claim 1, comprising means for appropriately controlling the publication of video data based on instructions received from the operating company, and for re-analyzing the data after the user has made improvements.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A