Content toxicity analysis method based on metadata
The metadata-based AI model for content toxicity analysis quickly and accurately assesses content toxicity by using intrinsic and extrinsic parameters, addressing inefficiencies in traditional methods and enhancing mental health content safety.
Patent Information
- Application Number
- PCT/KR2025/006995
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-06
- Filing Date
- 2025-05-23
- Publication Date
- 2026-02-12
AI Technical Summary
Existing content analysis methods fail to efficiently identify toxic elements that can exacerbate anxiety or depression, posing a risk to mental health, as they require extensive content analysis, which is time-consuming and inefficient.
A method utilizing metadata-based analysis through an AI model to extract keywords, calculate intrinsic parameter scores from comments, and determine content toxicity using extrinsic and intrinsic parameters, including relevance, fairness, objectivity, morality, and linguistics, without requiring full content analysis.
Enables rapid and accurate toxicity assessment of large volumes of content, reducing analysis time and improving efficiency while maintaining high accuracy in identifying potentially harmful content.
Smart Images

Figure KR2025006995_12022026_PF_FP_ABST
Abstract
Description
Metadata-based content toxicity analysis method
[0001] The present invention relates to a method for analyzing content toxicity, and more specifically, to an analysis method based on metadata of content.
[0002] The present invention is derived from research conducted as part of the Gyeonggi Regional Cooperation Research Center (GRRC) project of the Gyeonggi Regional Cooperation Research Center (Project Identification Number: GRRCAjou2023-B02, Ministry Name: Gyeonggi Province, Project Executing Organization Name: Ajou University Industry-Academic Cooperation Foundation, Research Period: 2024-07-01 ~ 2025-06-30, Project Name: Establishment of Patient Response Strategy Based on Clinical Big Data / Artificial Intelligence).
[0003] In addition, the present invention was derived from research conducted as part of the Korea Disease Control and Prevention Agency's Comprehensive Management of Health and Medical Bioresources (Project Unique Number: 2770000043, Project Number: 2024-ER0505-00, Project Management Organization Name: Korea Health Industry Development Institute, Research Project Name: Innovative Biobank Consortium for Chronic Cerebrovascular Disease Research, Project Performing Organization Name: Ajou University Industry-Academic Cooperation Foundation, Research Period: 2024.01.01~2024.12.31).
[0004] In addition, the present invention was derived from research conducted as part of the Ministry of Health and Welfare's research-oriented hospital promotion (Project unique number: 2460001063, Project number: RS-2022-KH130309, Project management organization name: Korea Health Industry Development Institute, Research project name: Super-gap SUPER*Senior PIMs (Project Innovation Management System) platform, Project performing organization name: Ajou University Industry-Academic Cooperation Foundation, Research period: 2024.01.01~2024.12.31).
[0005] In addition, the present invention is derived from research conducted as part of the Ministry of Science and ICT's bio and medical technology development (R&D) (Project unique number: 2710069629, Project number: RS-2021-NR056488, Project management organization name: National Research Foundation of Korea, Research project name: Aju SMART FARM Biocore Construction Project, Project performing organization name: Ajou University, Research period: 2024.01.01~2024.12.31).
[0006] Meanwhile, the Korean government, which provided the task, has no property interest in any aspect of the present invention.
[0007] Many therapeutic techniques involve providing content to individuals to promote mental health. Because content designed to promote mental health can directly impact individuals' emotions and mental states, toxicity analysis of such content is crucial. Toxic elements within content can exacerbate anxiety or depression, and even trigger trauma. Therefore, toxicity analysis technology is needed to identify potentially harmful elements within content to proactively prevent these risks.
[0008] One object of the present invention is to provide a method for analyzing content toxicity based on metadata.
[0009] According to one embodiment of the present invention, a method for analyzing content toxicity based on metadata can be provided.
[0010] Figure 1 is an environmental diagram of a content analysis system according to one embodiment.
[0011] Figure 2 is a flowchart of a content toxicity analysis method according to one embodiment.
[0012] Figure 3 is a flowchart of a method for obtaining scores for intrinsic parameters according to one embodiment.
[0013] Figure 4 is a flowchart of a method for determining the toxicity of content according to one embodiment.
[0014] Figure 5 is a flowchart of a toxicity analysis method according to another embodiment.
[0015] Figure 6 is a flowchart of a method for obtaining scores for intrinsic parameters according to another embodiment.
[0016] Figure 7 is a drawing for explaining the results produced by the content toxicity analysis method.
[0017] A content analysis method according to one embodiment is a content analysis method performed by at least one processor, the method including: inputting input data into an artificial intelligence model to obtain keywords; searching for content based on the keywords; selecting a first set of content using extrinsic parameters among a plurality of searched contents; obtaining comments related to the first set of contents; calculating a score for an intrinsic parameter based on the comments; and determining the toxicity of the first set of contents based on the score for the intrinsic parameter.
[0018] Here, the step of selecting a second set of contents based on the toxicity determined for each of the first set of contents may be further included; and the step of storing the second set of contents in a database.
[0019] Here, the extrinsic parameters may include at least one of a title, a description, a tag, an uploader, an upload date, views, recommendations, comments, length, image quality, subtitles, location information, license information, and a category.
[0020] Here, the intrinsic parameters may include at least one of relevance, fairness, objectivity, rightness, morality, expressiveness, and linguistics.
[0021] Here, the step of obtaining the keyword may be a step of obtaining keywords for each of the situation, symptom, and solution of the input data using the artificial intelligence model.
[0022] Here, the step of selecting the content of the first set may be a step of selecting at least one content among the content included in the searched content whose extrinsic parameters satisfy the selection criteria.
[0023] Here, the step of calculating a score for the intrinsic parameter includes the step of inputting the comment and the first prompt to the artificial intelligence model, and the step of obtaining a score for the relevance based on an output value of the artificial intelligence model for the first prompt, wherein the first prompt may include text related to whether the content of the content has a purpose of promoting mental health and text related to whether the content can affect the mental health of the viewer.
[0024] Here, the step of calculating a score for the intrinsic parameter includes the step of inputting the comment and the second prompt to the artificial intelligence model, and the step of obtaining a score for the fairness based on an output value of the artificial intelligence model for the second prompt, wherein the second prompt may include text related to whether the content is favorable to a specific object, text related to whether the content includes discrimination against a specific object, and text related to whether the content includes political content.
[0025] Here, the step of calculating a score for the intrinsic parameter includes the step of inputting the comment and the third prompt to the artificial intelligence model, and the step of obtaining a score for the objectivity based on an output value of the artificial intelligence model for the third prompt, wherein the third prompt may include text related to whether the content is objective and text related to whether the content is advertising content.
[0026] Here, the step of calculating a score for the intrinsic parameter includes the step of inputting the comment and the fourth prompt into the artificial intelligence model, and the step of obtaining a score for the rightness based on an output value of the artificial intelligence model for the fourth prompt, wherein the fourth prompt may include text related to whether the content of the content is defamatory, infringes on human rights, or infringes on privacy and portrait rights.
[0027] Here, the step of calculating a score for the intrinsic parameter includes the step of inputting the comment and the fifth prompt to the artificial intelligence model, and the step of obtaining a score for the morality based on an output value of the artificial intelligence model for the fifth prompt, wherein the fifth prompt may include text related to whether the content of the content violates social morality and text related to whether the content of the content causes harmful behavior.
[0028] Here, the step of calculating a score for the intrinsic parameter includes the step of inputting the comment and the sixth prompt to the artificial intelligence model, and the step of obtaining a score for the expressiveness based on an output value of the artificial intelligence model for the sixth prompt, wherein the sixth prompt may include text related to whether the content includes harmful content or expresses it in a beautified manner, text related to whether the content includes medical content with insufficient evidence, and text related to whether the content includes unscientific content.
[0029] Here, the step of calculating a score for the intrinsic parameter includes the step of inputting the comment and the seventh prompt to the artificial intelligence model, and the step of obtaining a score for the linguistic quality based on an output value of the artificial intelligence model for the seventh prompt, wherein the seventh prompt may include text related to whether the content includes profanity.
[0030] Here, the step of determining the toxicity may include a step of determining whether a score for each parameter included in the intrinsic parameters is greater than or equal to a first reference value, and a step of determining whether a score obtained by adding up the scores for each parameter included in the intrinsic parameters is greater than or equal to a second reference value.
[0031] Here, a computer program stored in a computer-readable recording medium may be provided to execute the above content analysis method.
[0032]
[0033] A computing device according to one embodiment includes a communication module; a memory; and at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory, wherein the at least one program may include instructions for inputting input data into an artificial intelligence model to obtain keywords, searching for content based on the keywords, selecting a first set of content using extrinsic parameters among a plurality of searched contents, obtaining comments related to the first set of contents, calculating a score for an intrinsic parameter based on the comments, and determining toxicity of the first set of contents based on the score for the intrinsic parameter.
[0034]
[0035] According to another embodiment, a content analysis method is provided, which is performed by at least one processor, and may include: obtaining vision data and sound data from content; calculating a score for an intrinsic parameter of the content based on the vision data and the sound data; and determining the toxicity of the content based on the score for the intrinsic parameter.
[0036] Here, the vision data may include at least one of a thumbnail, an object, an event, a color palette, and lighting conditions of the content.
[0037] Here, the sound data may include voice data and text data obtained based on the voice data.
[0038] Here, the intrinsic parameters may include at least one of relevance, fairness, objectivity, rightness, morality, expressiveness, and linguistics.
[0039] Here, the step of calculating a score for the intrinsic parameter includes the step of inputting the vision data, the sound data, and the first prompt into an artificial intelligence model, and the step of obtaining a score for the relevance based on an output value of the artificial intelligence model for the first prompt, wherein the first prompt may include text related to whether the content has a purpose of promoting mental health and text related to whether the content can affect the mental health of the viewer.
[0040] Here, the step of calculating a score for the intrinsic parameter includes the step of inputting the vision data, the sound data, and a second prompt into an artificial intelligence model, and the step of obtaining a score for the fairness based on an output value of the artificial intelligence model for the second prompt, wherein the second prompt may include text related to whether the content is favorable to a specific object, text related to whether the content includes discrimination against a specific object, and text related to whether the content includes political content.
[0041] Here, the step of calculating a score for the intrinsic parameter includes the step of inputting the vision data, the sound data, and a third prompt into an artificial intelligence model, and the step of obtaining a score for the objectivity based on an output value of the artificial intelligence model for the third prompt, wherein the third prompt may include text related to whether the content is objective and text related to whether the content is advertising content.
[0042] Here, the step of calculating a score for the intrinsic parameter includes the step of inputting the vision data, the sound data, and the fourth prompt into an artificial intelligence model, and the step of obtaining a score for the rightness based on an output value of the artificial intelligence model for the fourth prompt, wherein the fourth prompt may include text related to whether the content is defamatory, infringes on human rights, or infringes on privacy and portrait rights.
[0043] Here, the step of calculating a score for the intrinsic parameter includes the step of inputting the vision data, the sound data, and the fifth prompt into an artificial intelligence model, and the step of obtaining a score for the morality based on an output value of the artificial intelligence model for the fifth prompt, wherein the fifth prompt may include text related to whether the content violates social morality and text related to whether the content causes harmful behavior.
[0044] Here, the step of calculating a score for the intrinsic parameter includes the step of inputting the vision data, the sound data, and the sixth prompt into an artificial intelligence model, and the step of obtaining a score for the expressiveness based on an output value of the artificial intelligence model for the sixth prompt, wherein the sixth prompt may include text related to whether the content includes harmful content or expresses it in an embellished manner, text related to whether the content includes medical content with insufficient evidence, and text related to whether the content includes unscientific content.
[0045] Here, the step of calculating a score for the intrinsic parameter includes the step of inputting the vision data, the sound data, and the seventh prompt into an artificial intelligence model, and the step of obtaining a score for the linguistic quality based on an output value of the artificial intelligence model for the seventh prompt, wherein the seventh prompt may include text related to whether the content includes profanity.
[0046] Here, the step of determining the toxicity may include a step of determining whether a score for each parameter included in the intrinsic parameters is greater than or equal to a first reference value, and a step of determining whether a score obtained by adding up the scores for each parameter included in the intrinsic parameters is greater than or equal to a second reference value.
[0047] Here, the method further includes a step of selecting content based on whether an extrinsic parameter among a plurality of contents satisfies a selection criterion, wherein the content is content selected by an extrinsic parameter among a plurality of contents, and the extrinsic parameter may include at least one of a title, a description, a tag, an uploader, an upload date, a number of views, a number of recommendations, a number of comments, a length, a picture quality, subtitles, location information, license information, and a category.
[0048] Here, a step of obtaining comments related to the plurality of contents among the plurality of contents; a step of calculating a score for an intrinsic parameter for each of the plurality of contents based on the comments; and
[0049] The method further includes a step of selecting the content based on a score for the intrinsic parameter, wherein the intrinsic parameter may include at least one of relevance, fairness, objectivity, rightness, morality, expressiveness, and linguistics.
[0050] Here, a computer program stored on a computer-readable recording medium may be provided to execute a content analysis method.
[0051]
[0052] According to another embodiment, a computing device includes a communication module; a memory; and at least one processor connected to the memory and configured to execute at least one computer-readable program included in the memory, wherein the at least one program may include instructions for obtaining vision data and sound data from content, calculating a score for an intrinsic parameter of the content based on the vision data and the sound data, and determining toxicity of the content based on the score for the intrinsic parameter.
[0053] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the contents described in the attached drawings. However, the present invention is not limited or restricted by the exemplary embodiments. Unless otherwise defined, all terms (including technical and scientific terms) used in this specification are to be used with meanings that can be commonly understood by those of ordinary skill in the technical field to which this disclosure pertains. However, this may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc.
[0054] Additionally, terms defined in commonly used dictionaries should not be interpreted ideally or excessively unless explicitly and specifically defined otherwise. In certain cases, terms may be arbitrarily selected by the applicant, in which case their meanings will be described in detail in the relevant description. Therefore, the terms used in this disclosure should be defined based on their meaning and the overall content of this disclosure, rather than simply their names.
[0055] Throughout this specification, when a part is said to "include" a certain component, this does not mean that other components may be included, but rather that other components may be excluded, unless specifically stated otherwise. Furthermore, the singular forms used herein also include plural forms unless specifically stated otherwise. Furthermore, the expression "at least one of a, b, and / or c" used throughout this specification can encompass "a alone," "b alone," "c alone," "a and b," "a and c," "b and c," or "all of a, b, and c."
[0056] Meanwhile, terms such as "first and / or second" used in this specification may be used to describe various components, but are only used to distinguish one component from another and are not intended to be limited to the components referred to by those terms. For example, without departing from the scope of the present invention, the first component may be referred to as the second component, and the second component may also be referred to as the first component.
[0057] In addition, terms such as "unit" and "module" described in this specification mean a unit that processes at least one function or operation, which may be implemented by hardware or software or a combination of hardware and software. In addition, embodiments of the present disclosure in this specification may be represented by functional block configurations and various processing steps. These functional blocks may be implemented by various numbers of hardware or / and software configurations that execute specific functions. For example, embodiments of the present disclosure may employ direct circuit configurations such as memory, processing, logic, look-up tables, etc. that may execute various functions under the control of one or more microprocessors or other control devices.
[0058] In embodiments according to the present disclosure, functions related to artificial intelligence may be implemented through a processor and memory. In this case, the processor may be any one of a general-purpose processor such as a CPU (Center Processing Unit), an AP (Application Processor), a DSP (Digital Signal Processor), a graphics-only processor such as a GPU (Graphics Processing Unit), a VPU (Vision Processing Unit), and an AI-only processor such as an NPU (Neural Network Processing Unit). The processor may process input data according to predefined operation rules or AI models stored in the memory. Alternatively, if the processor is an AI-only processor, the AI-only processor may be designed with a hardware structure specialized for processing a specific AI model. In some embodiments according to the present disclosure, functions related to artificial intelligence may be implemented through a plurality of processors.
[0059] In embodiments of the present disclosure, predefined operating rules or artificial intelligence models may be configured to perform machine learning. Here, "configured to perform machine learning" means that the predefined operating rules or artificial intelligence models are trained using a learning algorithm and a plurality of learning data sets to perform a desired characteristic (or purpose). This learning may be performed within the device itself implementing the artificial intelligence according to the present disclosure, or may be performed through a separate server and / or system.
[0060] Artificial intelligence models can be implemented as neural networks (or artificial neural networks) and operate based on statistical learning algorithms that mimic biological neurons in machine learning and cognitive science. A neural network can refer to a general model in which artificial neurons (nodes) form a network by combining synapses, and through learning, the strength of the synaptic connections changes, thereby achieving problem-solving capabilities. A neural network can be composed of multiple neural network layers. For example, a neural network can include an input layer, a hidden layer, and an output layer. Each of the multiple neural network layers can include at least one node and at least one weight, and can perform neural network operations through operations between the computational results of the previous (precious) layer and the weights. At least one weight of the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, at least one weight can be updated during the learning process to reduce or minimize the loss or cost values obtained from the artificial intelligence model. Neural networks can infer the desired outcome from arbitrary input.
[0061] The learning methods of artificial intelligence models can be categorized into supervised learning, where input and output data are provided as training data, and the correct answer (output data) corresponding to the problem (input data) is determined, unsupervised learning, where only input data is provided without output data, and the correct answer (output data) corresponding to the problem (input data) is not determined, and reinforcement learning, where a reward is given whenever an action is taken in the current state, and learning progresses in the direction of maximizing this reward. Alternatively, they can be categorized according to the architecture, which is the structure of the learning model.
[0062] In an embodiment of the present disclosure, the artificial intelligence model is a CNN (Convolution Neural Network) such as GoogleNet, AlexNet, VGG Network, etc., R-CNN (Region with Convolution Neural Network), RPN (Region Proposal Network), RNN (Recurrent Neural Network), S-DNN (Stacking-based deep Neural Network), S-SDNN (State-Space Dynamic Neural Network), Deconvolution Network, DBN (Deep Belief Network), RBM (Restrcted Boltzman Machine), Fully Convolutional Network, LSTM (Long Short-Term Memory) Network, Classification Network, Generative Modeling, eXplainable AI, Continual AI, Representation Learning, AI for Material Design, BERT for natural language processing, SP-BERT, MRC / QA, Text Analysis, Dialog System, GPT-3, GPT-4, Visual Analytics for vision processing, Visual Understanding, Video Synthesis, ResNet for data intelligence, Anomaly Detection, Prediction, Time-Series Forecasting, Optimization, At least one of various artificial intelligence structures and algorithms, such as Recommendation, Data Creation, etc., may be used, and the above-described examples are merely listing examples of artificial intelligence structures and algorithms used according to embodiments of the present disclosure, and do not limit the artificial intelligence structures and algorithms used according to embodiments of the present disclosure.
[0063] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In describing the embodiments, descriptions of technical details that are well known in the technical field to which the present invention pertains and are not directly related to the present invention will be omitted. This is to convey the gist of the present invention more clearly without obscuring unnecessary explanation. For the same reason, some components in the accompanying drawings are exaggerated, omitted, or schematically depicted. Furthermore, the size of each component does not entirely reflect the actual size. Throughout this specification, the same reference numerals may refer to the same or corresponding components.
[0064]
[0065] Figure 1 is an environmental diagram of a content analysis system according to one embodiment.
[0066] Referring to FIG. 1, a content analysis system (100) according to one embodiment can communicate with a content database (10) and a user terminal (200) to transmit and receive data. The content analysis system (100) may be a computing device including at least one processor. In addition, the content analysis system (100) may include a communication module and memory.
[0067] The content analysis system (100) can obtain data from the content DB (10). At this time, the content DB (10) may be an external server. For example, the content DB (10) may be a search portal server or a content provision server (e.g., YouTube), but is not limited thereto. The content analysis system (100) can search and / or crawl content within the content DB (10) through search terms. In addition, the content analysis system (100) can obtain metadata, photo data, and video data contained within the content DB (10).
[0068] The content analysis system (100) can transmit and receive data by communicating with the user terminal (200). Specifically, the content analysis system (100) can obtain input data from the user terminal (200). For example, the content analysis system (100) can obtain text related to the user's current status from the user terminal (200). As a specific example, the content analysis system (100) can obtain text related to the user's concerns, such as "I'm anxious about finding a job," from the user terminal (200).
[0069] The content analysis system (100) can transmit content analysis results to the user terminal (200). Specifically, the content analysis system (100) can transmit information on recommended content based on input data and content analysis results to the user terminal (200). For example, if the content analysis system (100) obtains input data such as "I'm anxious about finding a job," it can select and transmit content with low toxicity and content that can improve the user's current mental health status to the user terminal (200).
[0070] In this way, the content analysis system (100) can analyze and / or evaluate content within the content DB (10) and store low-toxicity content in internal memory. In addition, the content analysis system (100) can transmit information on recommended content to the user terminal (200) based on data received from the user terminal (200).
[0071] To transmit information about recommended content to a user terminal (200), the content analysis system (100) may analyze and / or evaluate the toxicity of the content. In one embodiment, the content analysis system (100) may analyze and / or evaluate the toxicity of the content based on metadata related to the content, rather than the content itself. In another embodiment, the content analysis system (100) may analyze and / or evaluate the toxicity of the content using the content of the content. In yet another embodiment, the content analysis system (100) may analyze and / or evaluate the toxicity of the content using metadata and the content of the content.
[0072] The method by which the content analysis system (100) analyzes and / or evaluates the toxicity of content is described in detail below.
[0073]
[0074] Figure 2 is a flowchart of a content toxicity analysis method according to one embodiment.
[0075] Referring to FIG. 2, the content toxicity analysis method may include a step of obtaining keywords based on input data (S110), a step of searching content (S120), a step of selecting a first set of content (S130), a step of obtaining comments of the first set of content (S140), a step of calculating a score for an intrinsic parameter based on the comments (S150), and a step of determining the toxicity of the first set of content based on the intrinsic parameter (S160). Although FIG. 2 illustrates that steps S110 to S160 are performed sequentially, some steps may be omitted, some steps may be merged and performed simultaneously, or new steps may be added.
[0076] The step (S110) of obtaining keywords based on input data may be a step of inputting input data into an artificial intelligence model to obtain keywords. The input data may be data obtained from a user terminal (200). Specifically, the input data may be text data related to the user's status (or mental state) obtained from the user terminal (200). For example, the input data may be text data such as, but not limited to, "I feel like I'm being bullied," "Everyone else seems happy," "It's hard living with my parents," "I'm having a hard time at work," or "I cry for no reason after giving birth."
[0077] The AI model may be a generative AI model. Specifically, the AI model may be a large-scale language model. The AI model may be a model that outputs a text response in response to input data. The AI model may be contained in the memory of the content analysis system (100), but is not limited thereto. It may also be contained in an external server capable of communicating with the content analysis system (100).
[0078] In step S100, a processor included in the content analysis system (100, hereinafter referred to as the "system") may input input data into an artificial intelligence model. Furthermore, the processor may input instructions into the artificial intelligence model. For example, the processor may input the instruction prompt "Extract N keywords from the input data." Furthermore, for example, the processor may input the instruction prompt "Extract keywords for each of the situations, symptoms, and solutions from the input data."
[0079] Based on the input instructions, the AI model can output N keywords. For example, if the input data is "Everyone looks happy," the AI model can output the text "depression" for the situation, "depression, anxiety" for the symptom, and "nature, music, meditation, relaxation" for the solution. The processor can set the texts output by the AI model as keywords for the input data.
[0080] The step (S120) of searching for content may be a step of searching for content through the content DB (10) based on the keywords acquired in step S110. Specifically, the processor may search for content included in the content DB (10) using the keywords acquired in step S110 as search words. For example, the processor may search for content included in the content DB (10) using the keywords "depression, anxiety, meditation, relaxation" acquired in step S110 as search words.
[0081] Step S130 of selecting the first set of content may be a step in which the processor selects some content from among the plurality of contents searched in step S120 using extrinsic parameters. At this time, the extrinsic parameters may include at least one of the content's title, description, tags, uploader, upload date, number of views, number of recommendations, number of comments, length, image quality, subtitles, location information, license information, and category. Step S130 may be a step in which content is selected using metadata related to the posting of the content, rather than analyzing the content itself.
[0082] The processor may select a first set of content based on whether an extrinsic parameter among the plurality of contents satisfies the selection criteria. For example, the processor may select content with a view count of 10,000 or more from among the plurality of searched contents as the first set of content. Furthermore, the processor may select content with a comment count of 1,000 or more from among the plurality of searched contents as the first set of content. Furthermore, the processor may select content whose tags include search terms as the first set of content. Furthermore, the processor may select content whose upload date is within one year as the first set of content.
[0083] Step (S140) of obtaining comments on the first set of content may be a step in which the processor crawls the comments on each piece of content selected in step S130. For example, the processor may obtain comments on the first set of content. Additionally, for example, the processor may obtain descriptions of the first set of content.
[0084] Step S150 of calculating scores for intrinsic parameters based on comments may be a step in which the processor calculates scores for intrinsic parameters for each content in the first set based on the content-specific comments obtained in step S140. At this time, the intrinsic parameters may include at least one of relevance, fairness, objectivity, rightfulness, morality, expressiveness, and linguistics. Step S150, similar to step S130, may be a step in which scores for intrinsic parameters of the content are calculated using comments related to the content, rather than analyzing the content itself.
[0085] As in steps S130 and S150, the content toxicity analysis method according to one embodiment can quickly and efficiently determine the toxicity of content by analyzing the content's metadata rather than analyzing the content itself. Therefore, unlike conventional methods that analyze all of the content, the toxicity analysis method using metadata can analyze a large amount of content in a short period of time and achieve high accuracy. A detailed description of step S150 is provided below with reference to FIG. 3.
[0086] The step (S160) of determining the toxicity of the first set of content based on the intrinsic parameters may be a step in which the processor determines the toxicity of the first set of content based on the score for the intrinsic parameters calculated in step S150. A detailed description of step S160 is provided below with reference to FIG. 4.
[0087]
[0088] Figure 3 is a flowchart of a method for obtaining scores for intrinsic parameters according to one embodiment.
[0089] Referring to FIG. 3, a method for obtaining scores for intrinsic parameters according to one embodiment may include a step of inputting comments and prompts (S151) and a step of obtaining scores for each intrinsic parameter (S152). While FIG. 3 illustrates steps S151 to S152 being performed sequentially, this is not limited thereto and new steps may be added.
[0090] The step of inputting comments and prompts (S151) may be a step in which the processor inputs the comments and prompts obtained in step S140 into an artificial intelligence model. At this time, the artificial intelligence model may be the same as the artificial intelligence model used to obtain keywords in step S110, but is not limited thereto.
[0091] Additionally, in step S151, the processor may additionally input an instruction prompt to the AI model to obtain a score for the intrinsic parameters. For example, the instruction prompt may be text such as, but not limited to, "Give me a score between -5 and 5 for the input item. -5 is very bad, 5 is very good, and 0 is unknown."
[0092] Specifically, to obtain a score related to relevance among the intrinsic parameters, the processor may input a comment and a first prompt to the AI model. The processor may obtain a score related to relevance among the intrinsic parameters based on the output of the AI model for the first prompt.
[0093] At this time, the first prompt may include text regarding whether the content is intended to promote mental health and whether the content may impact the viewer's mental health. For example, the first prompt may include, but is not limited to, text such as, "Is the content of the comment consistent with the purpose of promoting and treating mental health?" or "Can this content impact the viewer's emotions, cognition, and behavior regarding mental health?"
[0094] Additionally, to obtain a score related to fairness among the intrinsic parameters, the processor may input comments and a second prompt to the AI model. The processor may obtain a score for fairness among the intrinsic parameters based on the AI model's output in response to the second prompt.
[0095] Here, the second prompt may include text regarding whether the content favors a specific entity, whether the content includes discrimination against a specific entity, and whether the content includes political content. For example, the second prompt may include, but is not limited to, text such as, "When distinguishing between facts and commentary, the content presents the truth in a balanced manner without distorting or biasing social issues," "The content must not be designed to favor a specific group or stakeholder or mislead the public," or "The content must not discriminate based on various criteria such as gender, age, occupation, religion, or race, or undermine fairness and equity in elections and political issues."
[0096] Additionally, to obtain a score related to objectivity among the intrinsic parameters, the processor may input comments and a third prompt to the AI model. The processor may obtain a score for objectivity among the intrinsic parameters based on the AI model's output for the third prompt.
[0097] At this time, the third prompt may include text regarding the objectiveness of the content and whether the content is advertising. For example, the third prompt may include, but is not limited to, text such as, "The content should accurately and objectively address facts, avoid broadcasting information of unclear origin, and clearly cite the sources of materials used." or "The content should not imply conflicts of interest, such as product advertising or political affiliation."
[0098] Additionally, to obtain a score related to the "infringement of rights" intrinsic parameter, the processor may input a comment and a fourth prompt to the AI model. The processor may obtain a score related to the "infringement of rights" intrinsic parameter based on the AI model's output for the fourth prompt.
[0099] At this time, the fourth prompt may include text regarding whether the content violates defamation, human rights, privacy, and portrait rights. For example, the fourth prompt may include, but is not limited to, text such as, "Prohibits invasion of privacy, defamation, and human rights violations" or "Does not include private conversations or portraits of others without their consent."
[0100] Additionally, to obtain a score related to the ethical level among the intrinsic parameters, the processor can input comments and a fifth prompt to the AI model. The processor can obtain a score for the ethical level among the intrinsic parameters based on the AI model's output for the fifth prompt.
[0101] At this time, the fifth prompt may include text regarding whether the content violates social morality and whether the content encourages harmful behavior. For example, the fifth prompt may include, but is not limited to, text such as, "Are public morality, social and family values, and respect for life being destroyed?", "Gender equality, cultural diversity, and religious freedom must be respected.", or "Actions that glorify or promote drinking, smoking, gambling, extravagance, and wastefulness should not be permitted."
[0102] Additionally, to obtain a score related to expressiveness (Methods and Expressions) among the intrinsic parameters, the processor may input comments and a sixth prompt to the AI model. The processor may obtain a score for expressiveness among the intrinsic parameters based on the AI model's output for the sixth prompt.
[0103] Here, the sixth prompt may include text regarding whether the content contains or glorifies harmful content, whether the content contains medical claims with insufficient evidence, or whether the content contains unscientific content. For example, the sixth prompt may include, but is not limited to, text such as, "Content should not specifically depict or glorify sexual, violent, or criminal content, and if an event is reenacted, it must be clearly disclosed," "Content should not contain medical or health-related content with insufficient scientific evidence," "Content should not promote superstitions or unscientific attitudes," or "Content should not directly depict scenes or methods of suicide or glorify suicide as a solution to life's suffering."
[0104] Additionally, to obtain a score related to language among the intrinsic parameters, the processor may input a comment and a seventh prompt to the AI model. The processor may obtain a score for language among the intrinsic parameters based on the AI model's output for the seventh prompt.
[0105] At this time, the seventh prompt may include text regarding whether the content contains profanity. For example, the seventh prompt may include, but is not limited to, text such as, "Do not use accents, tones, slang, vulgar expressions, or profanity that disrupt proper language habits."
[0106] The step (S152) of obtaining a score for each intrinsic parameter may be a step of obtaining a value output by an artificial intelligence model based on the comments, prompts, and instructions input in step S151. The artificial intelligence model may output a score for each of relevance, fairness, objectivity, rightfulness, morality, expressiveness, and linguistics. However, the present invention is not limited thereto, and the artificial intelligence model may also output a score for a single intrinsic parameter that combines relevance, fairness, objectivity, rightfulness, morality, expressiveness, and linguistics, based on the input instruction prompt.
[0107]
[0108] Figure 4 is a flowchart of a method for determining the toxicity of content according to one embodiment.
[0109] Referring to FIG. 4, a method for determining the toxicity of content according to one embodiment may include a step (S161) of determining whether a score for each intrinsic parameter is greater than or equal to a first reference value, a step (S162) of determining whether a sum score of the intrinsic parameters is greater than or equal to a second reference value, and a step (S163) of determining the toxicity of the content. Although FIG. 4 illustrates that steps S161 to S163 are performed sequentially, this is not limited thereto, and the order of the steps may be changed, some steps may be omitted, some steps may be merged and performed simultaneously, or new steps may be added. For example, the order of steps S161 and S162 may be changed.
[0110] The step (S161) of determining whether the score for each intrinsic parameter is greater than or equal to a first reference value may be a step in which the processor determines whether the score for each intrinsic parameter obtained through the output value of the artificial intelligence model is greater than or equal to a certain value. Specifically, the processor may verify the score for each intrinsic parameter, including relevance, fairness, objectivity, rightness, morality, expressiveness, and linguistics. The processor may determine whether each score is greater than or equal to the first reference value.
[0111] If the processor identifies an intrinsic parameter with a score lower than the first reference value, the processor may determine that the content is toxic by going through step S163 without going through step S162. Alternatively, if there are N or more intrinsic parameters with scores lower than the first reference value (N is a natural number greater than or equal to 2), the processor may determine that the content is toxic by going through step S163 without going through step S162. However, the present invention is not limited thereto, and even if there are intrinsic parameters with scores lower than the first reference value, the processor may perform step S162.
[0112] The step (S162) of determining whether the sum score of the intrinsic parameters is greater than or equal to the second reference value may include a step in which the processor calculates scores for each of the intrinsic parameters, including relevance, fairness, objectivity, rightness, morality, expressiveness, and linguistics. The range of the sum score may be -35 to 35, but is not limited thereto.
[0113] The step (S163) of determining the toxicity of content may be a step in which the processor determines the toxicity of the content based on the determination results of steps S161 and S162. In step S163, the processor may determine whether the content is toxic or may calculate a score for the toxicity of the content.
[0114] For example, if one or more negative results are found among the results determined in step S161, the processor may determine that the content is toxic. Furthermore, for example, if the score calculated in step S162 is less than the second reference value, the processor may determine that the content is toxic. Furthermore, for example, the processor may determine the level (or class) of toxicity of the content (e.g., non-toxic / weakly toxic / highly toxic, etc.) based on the total score calculated in step S162.
[0115] After step S163, the processor may select a second set of content from the first set of content based on the toxicity of the content. The processor may store the selected second set of content in the memory (database) of the content analysis system (100).
[0116] For example, the processor may select content determined to be non-toxic in step S163 as a second set of content. Furthermore, for example, the processor may select content whose toxicity score calculated in step S163 is below a third reference value as a second set of content. Furthermore, for example, the processor may select content whose toxicity level (or class) calculated in step S163 is a specific level (or class) as a second set of content.
[0117]
[0118] Figure 5 is a flowchart of a toxicity analysis method according to another embodiment.
[0119] Referring to FIG. 5, a toxicity analysis method according to another embodiment may include a step of acquiring vision data and sound data (S210), a step of calculating scores for intrinsic parameters based on the vision data and sound data (S220), and a step of determining the toxicity of content based on the intrinsic parameters (S230). Although FIG. 5 illustrates that steps S210 to S230 are performed sequentially, this is not limited to the sequential execution of some steps, some steps may be omitted, some steps may be merged and performed simultaneously, or new steps may be added.
[0120] The step of acquiring vision data and sound data (S210) may be a step of acquiring vision data related to images and sound data related to audio from content. At this time, the vision data may include at least one of a thumbnail, object, event, color palette, and lighting conditions of the content. Additionally, the sound data may include audio data and text data acquired based on the audio data.
[0121] The processor can acquire vision data by dividing content into frames and extracting individual images. Additionally, the processor can perform image preprocessing on each extracted frame and adjust the resolution by removing noise. Furthermore, the processor can apply an object detection algorithm to the preprocessed image to identify and locate specific objects or people within the frame. Furthermore, the processor can track the movement of the identified objects or people and analyze changes in position over time. Furthermore, the processor can detect scene transitions by analyzing changes between frames. Furthermore, the processor can analyze the colors of the frames to extract a color palette and lighting conditions.
[0122] The processor can obtain sound data by isolating and extracting audio tracks from content. Additionally, the processor can digitize the extracted audio signals and adjust the sampling rate. Furthermore, the processor can divide the audio signals into frames and perform analysis on each frame. Furthermore, the processor can apply a Fourier transform to each audio frame to extract frequency components and visualize them as a spectrogram. Furthermore, the processor can perform various audio analyses, such as voice recognition, emotion analysis, and background noise detection, based on the spectrogram.
[0123] The processor can obtain text data based on speech data extracted from content. Specifically, the processor can extract acoustic features from the speech data and input the acoustic features into an acoustic model to convert them into phonemes. In the process of combining phoneme sequences into words, the processor uses a language model to select contextually appropriate words and combines the selected words into sentences, ultimately generating text data. A detailed description of other speech recognition technologies is omitted.
[0124] The step (S220) of calculating scores for intrinsic parameters based on vision data and sound data may be a step in which the processor inputs the data acquired in step S210 into an artificial intelligence model, thereby obtaining scores for intrinsic parameters using the output values of the artificial intelligence model. A detailed description of step S220 is provided below with reference to FIG. 6.
[0125] The step (S230) of determining the toxicity of content based on intrinsic parameters may be similar to step S160 of FIG. 2 . The method for determining content toxicity of FIG. 2 determines toxicity using metadata rather than the content itself, while the method for determining content toxicity of FIG. 5 determines toxicity using the content itself. The two methods differ in certain aspects. Specifically, FIG. 2 uses only metadata, while FIG. 5 uses vision data and sound data related to the content itself. However, the method for determining content toxicity using intrinsic parameters may be similar.
[0126] Step S230 for determining the toxicity of content may be a step in which the processor determines the toxicity of the content based on the scores for the intrinsic parameters of step S220. The processor may determine the presence or absence of toxicity of the content in step S230, or may calculate a score for the toxicity of the content. The following description will be given as an example where step S230 includes steps S161 to S163 of FIG. 4.
[0127] For example, if one or more negative results are found among the results determined in step S161, the processor may determine that the content is toxic. Furthermore, for example, if the score calculated in step S162 is less than the second reference value, the processor may determine that the content is toxic. Furthermore, for example, the processor may determine the level (or class) of toxicity of the content (e.g., non-toxic / weakly toxic / highly toxic, etc.) based on the total score calculated in step S162.
[0128] After step S163, the processor may select some content based on its toxicity. The processor may store the selected content in the memory (database) of the content analysis system (100).
[0129] For example, the processor may select content determined to be non-toxic in step S163 and store it in memory. Furthermore, for example, the processor may select content whose toxicity score calculated in step S163 is below a third reference value and store it in memory. Furthermore, for example, the processor may select content whose toxicity level (or class) calculated in step S163 is a specific level (or class) and store it in memory.
[0130]
[0131] Figure 6 is a flowchart of a method for obtaining scores for intrinsic parameters according to another embodiment.
[0132] Referring to FIG. 6, a method for obtaining scores for intrinsic parameters according to another embodiment may include a step (S221) of inputting vision data, sound data, and a prompt, and a step (S222) of obtaining scores for each intrinsic parameter. FIG. 6 illustrates that steps S221 and S222 are performed sequentially, but this is not limited thereto, and new steps may be added.
[0133] The step of inputting vision data, sound data, and prompts (S221) may be a step in which the processor inputs the data and prompts acquired in step S210 into the artificial intelligence model.
[0134] Additionally, in step S221, the processor may additionally input an instruction prompt to the AI model to obtain a score for the intrinsic parameters. For example, the instruction prompt may be text such as, but not limited to, "Give me a score between -5 and 5 for the input item. -5 is very bad, 5 is very good, and 0 is unknown."
[0135] Specifically, to obtain scores related to relevance, fairness, objectivity, rightfulness, morality, expressiveness, and linguistics among the intrinsic parameters, the processor may input comments and the first to seventh prompts to the artificial intelligence model. The processor may obtain scores related to relevance, fairness, objectivity, rightfulness, morality, expressiveness, and linguistics among the intrinsic parameters based on the output values of the artificial intelligence model for the first to seventh prompts. Since the contents of the first to seventh prompts overlap with those of FIG. 3, a detailed description thereof will be omitted.
[0136] Step S222 of obtaining a score for each intrinsic parameter may be similar to step S152 of FIG. 3. Step S222 may be a step of obtaining a value output by an artificial intelligence model based on the comments, prompts, and instructions input in step S221. The artificial intelligence model may output scores for each of relevance, fairness, objectivity, rightfulness, morality, expressiveness, and linguistics. However, the present invention is not limited thereto, and the artificial intelligence model may also output a score for a single intrinsic parameter that synthesizes relevance, fairness, objectivity, rightfulness, morality, expressiveness, and linguistics based on the input instruction prompt.
[0137]
[0138] According to another embodiment, a content analysis method may be a hybrid method of the method of FIG. 2 and the method of FIG. 5 . For example, the processor may first evaluate the toxicity of content based on metadata using the method of FIG. 2 to initially select content. Thereafter, the processor may secondarily evaluate the toxicity of content based on vision data and sound data related to the content itself using the method of FIG. 5 to secondarily select content. Accordingly, the content analysis system (100) may ultimately store the secondarily selected content in memory. The content analysis system (100) may transmit information about the secondarily selected content stored in the memory to the user terminal (200).
[0139] Specifically, the content for acquiring vision data and sound data in step S210 of FIG. 5 may be the content selected in FIG. 2. That is, step S210 of FIG. 5 may be a step for acquiring vision data and sound data from the selected content. At this time, the selected content may be content selected based on whether an extrinsic parameter satisfies a selection criterion and / or content selected based on a score of an intrinsic parameter.
[0140]
[0141] Figure 7 is a drawing for explaining the results produced by the content toxicity analysis method.
[0142] Referring to Figure 7, the results of the content toxicity analysis can be confirmed. The processor can obtain scores for each intrinsic parameter of the content using an AI model. Furthermore, the processor can obtain a toxicity assessment of the content using the AI model.
[0143] The processor can use the results obtained through the artificial intelligence model to provide customized content to the user terminal (200). For example, the processor can analyze the toxicity assessment results of FIG. 7 to select users for whom the content will be effective, and provide information about the content to the selected users.
[0144] Additionally, the processor can use tags to group content based on the toxicity assessment of each content. For example, the processor can create various content groups, such as a first content group related to depression treatment, a second content group related to eating disorder treatment, and a third content group related to nicotine addiction treatment. In other words, the processor can group and manage content based on the toxicity assessment.
[0145]
[0146] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination. The program commands recorded on the medium may be those specially designed and configured for the embodiment or may be those known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of the program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operations of the embodiment, and vice versa.
[0147] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
[0148] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.
Claims
1. In a content analysis method performed by at least one processor, A step of obtaining keywords by inputting input data into an artificial intelligence model; A step of searching content based on the above keywords; A step of selecting a first set of contents from among multiple searched contents using extrinsic parameters; A step of obtaining comments related to the content of the first set; A step of calculating a score for an intrinsic parameter based on the above comments; and A step of determining the toxicity of the first set of contents based on the score for the intrinsic parameter. Content analysis methods.
2. In paragraph 1, A step of selecting a second set of contents based on the toxicity determined for each of the first set of contents; and Further comprising the step of storing the contents of the second set in a database. Content analysis methods.
3. In paragraph 1, The above extrinsic parameters include at least one of title, description, tags, uploader, upload date, views, recommendations, comments, length, quality, subtitles, location information, license information, and category. Content analysis methods.
4. In paragraph 1, The above intrinsic parameters include at least one of relevance, fairness, objectivity, rightness, morality, expressiveness and linguistics. Content analysis methods.
5. In paragraph 1, The steps for obtaining the above keywords are: A step of obtaining keywords for each situation, symptom, and solution of the input data using the above artificial intelligence model. Content analysis methods.
6. In paragraph 1, The step of selecting the content of the first set above is: A step of selecting at least one content among the contents included in the searched contents whose external parameters meet the selection criteria. Content analysis methods.
7. In paragraph 4, The step of calculating the score for the above intrinsic parameters is: a step of inputting the comment and the first prompt into the artificial intelligence model, and A step of obtaining a score for the relevance based on the output value of the artificial intelligence model for the first prompt, The first prompt above includes text related to whether the content is intended to promote mental health and text related to whether the content may affect the mental health of the viewer. Content analysis methods.
8. In paragraph 4, The step of calculating the score for the above intrinsic parameters is: a step of inputting the comment and the second prompt into the artificial intelligence model, and A step of obtaining a score for the fairness based on the output value of the artificial intelligence model for the second prompt, The second prompt above includes text related to whether the content is favorable to a specific object, text related to whether the content includes discrimination against a specific object, and text related to whether the content includes political content. Content analysis methods.
9. In paragraph 4, The step of calculating the score for the above intrinsic parameters is: a step of inputting the comment and the third prompt into the artificial intelligence model, and A step of obtaining a score for the objectivity based on the output value of the artificial intelligence model for the third prompt is included, The third prompt above includes text regarding whether the content is objective and text regarding whether the content is advertising. Content analysis methods.
10. In paragraph 4, The step of calculating the score for the above intrinsic parameters is: a step of inputting the comment and the fourth prompt into the artificial intelligence model, and A step of obtaining a score for the rightness based on the output value of the artificial intelligence model for the fourth prompt is included, The fourth prompt above contains text regarding whether the content is defamatory, infringing on human rights, or violating privacy and portrait rights. Content analysis methods.
11. In paragraph 4, The step of calculating the score for the above intrinsic parameters is: A step of inputting the comment and the fifth prompt into the artificial intelligence model, and A step of obtaining a score for the morality based on the output value of the artificial intelligence model for the fifth prompt is included. The fifth prompt above includes text related to whether the content is against social morality and text related to whether the content causes harmful behavior. Content analysis methods.
12. In paragraph 4, The step of calculating the score for the above intrinsic parameters is: A step of inputting the comment and the sixth prompt into the artificial intelligence model, and A step of obtaining a score for the expressiveness based on the output value of the artificial intelligence model for the sixth prompt, The sixth prompt above includes text related to whether the content contains or embellishes harmful content, text related to whether the content contains medical claims that lack evidence, and text related to whether the content contains unscientific claims. Content analysis methods.
13. In paragraph 4, The step of calculating the score for the above intrinsic parameters is: A step of inputting the comment and the seventh prompt into the artificial intelligence model, and A step of obtaining a score for the linguistic quality based on the output value of the artificial intelligence model for the seventh prompt is included. The seventh prompt above contains text regarding whether the content contains profanity. Content analysis methods.
14. In paragraph 1, The steps for judging the above toxicity are: A step of determining whether the score for each parameter included in the above intrinsic parameters is greater than or equal to the first reference value, and A step of determining whether the sum of the scores for each parameter included in the above intrinsic parameters is greater than or equal to the second reference value. Content analysis methods.
15. A non-transitory computer-readable recording medium having recorded thereon a program for executing the content analysis method described in Article 1.
16. Communication module; memory; and At least one processor connected to said memory and configured to execute at least one computer-readable program contained in said memory, At least one of the above programs, Input data is input into an artificial intelligence model to obtain keywords, Search for content based on the above keywords, Select the first set of contents from among the searched multiple contents using extrinsic parameters, Obtain comments related to the content of the first set above, Based on the above comments, we calculate the scores for the intrinsic parameters, comprising instructions for determining the toxicity of the first set of contents based on a score for the intrinsic parameter; Computing device.
Citation Information
Patent Citations
System and method for multi-stage filtering of malicious videos in video distribution environment
KR1020090006397A
Media contents sharing system and method using media contents filtering
KR1020140119229A
Apparatus and Method of Video Contents Recommendation based on Emotion Ontology
KR1020160143411A
Separation device for seaweed stem and lobe
KR102770265B1