Context-based multimedia content moderation
The dynamic multimedia content moderation system addresses the lack of contextual understanding in existing systems by using region-specific profanity and user profiles, improving content filtering accuracy and personalization through adaptive AI-based moderation.
Patent Information
- Application Number
- PCT/IB2025/056644
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-04
- Filing Date
- 2025-06-30
- Publication Date
- 2026-01-08
AI Technical Summary
Existing multimedia content moderation systems lack contextual understanding and fail to account for regional variations in language and cultural norms, leading to suboptimal user experiences due to over-blocking or under-filtering of inappropriate content.
A dynamic content moderation system that utilizes context information, region-specific profanity data, and user profile information to adaptively filter content, incorporating advanced technologies like AI models and real-time sensor data for personalized moderation.
Enhances content moderation accuracy, adapts to regional and cultural sensitivities, and provides personalized experiences by dynamically updating moderation based on real-time user preferences and environmental changes.
Smart Images

Figure IB2025056644_08012026_PF_FP_ABST
Abstract
Description
CONTEXT-BASED MULTIMEDIA CONTENT MODERATIONCROSS-REFERENCE TO RELATED APPLICATIONS / INCORPORATION BY REFERENCE
[0001] This application claims priority to Indian Provisional Application No. IN 202411051426, filed July 4, 2024, which is hereby incorporated by reference in its entirety.FIELD
[0002] Various embodiments of the disclosure relate to multimedia content moderation systems. More specifically, various embodiments of the disclosure relate to an electronic device and a method for dynamic content moderation based on content context.BACKGROUND
[0003] Multimedia content moderation systems have evolved to address concerns related to inappropriate or offensive content in various forms of media. These systems aim to filter and control the presentation of content based on user preferences, age restrictions, and cultural sensitivities. Advancements in this domain have led to the development of automated content analysis techniques and user-specific filtering mechanisms.
[0004] Typical content moderation solutions employ methods such as keyword filtering, image recognition, and audio analysis to identify potentially unsuitable content. Some systems utilize pre-defined rating systems or content tags to categorize media segments. However, these approaches often lack contextual understanding and may not account for regional variations in language and cultural norms. Additionally, existing solutions may struggle to adapt to individual user preferences or real-time changes in viewingenvironments. These limitations can result in over-blocking of harmless content or underfiltering of inappropriate material, leading to suboptimal user experiences and potential exposure to unsuitable content.
[0005] Limitations and disadvantages of conventional and traditional approaches will become apparent to one of skill in the art, through comparison of described systems with some aspects of the present disclosure, as set forth in the remainder of the present application and with reference to the drawings.SUMMARY
[0006] An electronic device and method for dynamic content moderation based on content context, region-specific profanity, and user profile information is provided substantially as shown in, and / or described in connection with, at least one of the figures, as set forth more completely in the claims.
[0007] These and other features and advantages of the present disclosure may be appreciated from a review of the following detailed description of the present disclosure, along with the accompanying figures in which like reference numerals refer to like parts throughout.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] FIG. 1 is a block diagram that illustrates an exemplary network environment for dynamic content moderation based on content context, region-specific profanity, and user profile information, in accordance with an embodiment of the disclosure.
[0009] FIG. 2A is a block diagram that illustrates an exemplary first electronic device of FIG. 1 , in accordance with an embodiment of the disclosure.
[0010] FIG. 2B is a block diagram that illustrates an exemplary second electronicdevice of FIG. 1 , in accordance with an embodiment of the disclosure.
[0011] FIG. 3 is a diagram that illustrates an exemplary scenario of a content moderation system, in accordance with at least one embodiment of the disclosure.
[0012] FIG. 4 is a diagram that illustrates an exemplary scenario of a content moderation system, in accordance with an embodiment of the disclosure.
[0013] FIG. 5 is a flowchart illustrating an example method for multimedia content moderation based on content context, in accordance with an embodiment of the disclosure.DETAILED DESCRIPTION.
[0014] The present disclosure relates to multimedia content moderation based on content context, region-specific profanity, and user profile. The following described implementation may be found in an electronic device and method that may be configured to dynamically moderate content based on content context, region-specific profanity, and user profile information. Exemplary aspects of the present disclosure may provide an electronic device that may be configured to receive multimedia content and content metadata associated with a plurality of segments of the multimedia content from a second electronic device. The electronic device may be further configured to determine context information associated with each segment of the plurality of segments, based on the content metadata. The electronic device may be further configured to receive regionspecific profanity information associated with the multimedia content from the second electronic device. The electronic device may be further configured to receive, from a set of sensors, sensor data associated with a user present in a predetermined proximity of the first electronic device. The electronic device may be further configured to determineprofile information associated with the user based on the sensor data. The electronic device may be further configured to apply a content moderation model on the plurality of segments, based on the context information, the region-specific profanity information, and the profile information. The electronic device may be further configured to select, from the plurality of segments, a set of unsuitable segments for the user, based on the application of the content moderation model. The electronic device may be further configured to determine a moderated media content based on the multimedia content and the set of unsuitable segments. The electronic device may be further configured to control a user device to render the moderated media content.
[0015] Existing content moderation systems often face challenges in provision of contextually appropriate content based on personalized content filtering. Current solutions typically rely on pre-defined rating systems or content tags, which may not account for regional variations in language and cultural norms. These systems often lack the ability to adapt to individual user preferences or real-time changes in viewing environments. As a result, conventional approaches may lead to over-blocking of harmless content or under-filtering of inappropriate material, which may result in suboptimal user experiences and potential exposure to unsuitable content.
[0016] The disclosed content moderation technique utilizes a combination of advanced technologies to provide a more nuanced and adaptive approach to content filtering. By leveraging context information, region-specific profanity data, and user profile information, the disclosed technique may make more informed decisions about content suitability. Unlike traditional systems, the disclosed approach may dynamically adjust moderation parameters based on the specific characteristics of the content, regionalnorms, and individual user profiles. The disclosed technique may offer improved accuracy in identification of unsuitable content, better adaptation to regional and cultural sensitivities, and enhanced personalization of content moderation based on user profiles. Additionally, the disclosed technique may have an ability to dynamically update the moderation process based on real-time sensor data, which may allow a more responsive and context-aware content filtering experience.
[0017] FIG. 1 is a block diagram that illustrates an exemplary network environment for multimedia content moderation based on content context, region-specific profanity, and user profile, in accordance with an embodiment of the disclosure. With reference to FIG. 1 , there is shown a network environment 100. The network environment 100 may include a first electronic device 102, a second electronic device 104, a server 106, a database 108, a communication network 112, and a user device 114. In an embodiment, the first electronic device 102 may include a content moderation model 102A, and a set of sensors 102B. Further, the second electronic device 104 may include a Vision Foundation Model (VFM) 104A, a Neural Language Model (NLM) 104B, and an Audio Language Model (ALM) 104C. The first electronic device 102 and the second electronic device 104 may be connected to the server 106 through the communication network 112. The server 106 may be coupled to the database 108, which may store the multimedia content 110A and a content metadata 110B. The user device 114 may also be connected to the communication network 112, which may allow communication of the user device 114 with the first electronic device 102, the second electronic device 104, and access to the server 106. In an embodiment, the user device 114 may be associated with at least one user, such as a user 114A.
[0018] The first electronic device 102 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive the multimedia content 110A and the content metadata 110B associated with a plurality of segments of the multimedia content 110A from the second electronic device 104. The first electronic device 102 may determine context information associated with each segment of the plurality of segments, based on the content metadata 110B. The first electronic device 102 may receive regionspecific profanity information associated with the multimedia content 110A from the second electronic device 104. The first electronic device 102 may further receive, from the set of sensors 102B, sensor data associated with the user 114A, who may be present in a predetermined proximity of the first electronic device 102. The first electronic device 102 may determine profile information associated with the user 114A based on the sensor data. The first electronic device 102 may apply a content moderation model 102A on the plurality of segments, based on the context information, the region-specific profanity information, and the profile information. The first electronic device 102 may select, from the plurality of segments, a set of unsuitable segments for the user, based on the application of the content moderation model 102A. The first electronic device 102 may determine a moderated media content based on the multimedia content 110A and the set of unsuitable segments and may control the user device 114 to render the moderated media content. Examples of the first electronic device 102 may include, but are not limited to, a computing device, a server, a network provider, a base station, a router, a smartphone, a cellular phone, a mobile phone, a gaming device, a mainframe machine, a computer workstation, a consumer electronic (CE) device, a television, a set-top box, a broadcaster, a base station device, and / or the like.
[0019] In an embodiment, the first electronic device 102 may further be configured to determine an age of the user 114A based on the sensor data and determine the profile information based on the age of the user. The first electronic device 102 may further be configured to detect a change in the sensor data to update the profile information based on the change and dynamically update the moderated media content based on the updated profile information.
[0020] In an embodiment, the first electronic device 102 may further be configured to generate a content recommendation based on the profile information and control the user device 114 to render the content recommendation. The first electronic device 102 may be further configured to determine a geographic location of the first electronic device 102 or the user device 114 and receive the region-specific profanity information based on the geographic location.
[0021] In an embodiment, the first electronic device 102 may mask or replace the set of unsuitable segments in the moderated media content. Further, the first electronic device 102 may receive user input (of the user 114A) to modify the content moderation model 102A and update the content moderation model 102A based on the user input. The first electronic device 102 may generate a content moderation report based on the set of unsuitable segments and transmit the content moderation report to the second electronic device 104. As used herein, the term “unsuitable segments” may portions / sections of the multimedia content 110A that may be identified as inappropriate, irrelevant, or inconsistent with predefined moderation criteria or user preferences.Examples of unsuitable segments may include, but not limited to inappropriate content such as, content with offensive language, explicit visuals, sensitive themes, or the likes,irrelevant information (such as, sections unrelated to the context or purpose of the media) and distracting elements (such as, noisy or redundant portions that do not add value to the intended user experience).
[0022] The content moderation model 102A may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive the receive multimedia content 110A and content metadata 11 OB, the region-specific profanity information, and the profile information associated with the user 114A. The content moderation model 102A may detect the context information associated with each segment of the plurality of segments, based on the content metadata 11 OB. The content moderation model 102A may further be configured to determine, from the plurality of segments, the set of unsuitable segments for the user 114A. The content moderation model 102A may be updated based on the user input received by the first electronic device 102.
[0023] In an exemplary embodiment, the content moderation model 102A may be configured to manage and regulate the multimedia content 110A on platforms such as a Video-On-Demand (VOD) server, an Over-The-Top (OTT) platform, a Television (TV) broadcasting platform. The content moderation model 102A may ensure that the multimedia content 110A aligns with the platform’s profanity information, rules and ethical standards and also protect the user 114A from the set of unsuitable segments of the multimedia content 110A. The content moderation model 102A may use a combination of automated systems, such as an artificial Intelligence (Al), to detect profanity information, and human moderators (such as a sensor board) to review the set of unsuitable segments for context and accuracy. The content moderation model 102A may include community guidelines that may define acceptable behavior and a set of suitablesegments from the plurality of segments of the multimedia content 110A. Also, the content moderation model 102A may be configured to maintain a safe online environment based on a balance between free expression with restrictions, management of cultural differences, and adaptation to evolving harmful content trends.
[0024] In an embodiment, the content moderation model 102A may be at least one of: a pre-moderation model, a post-moderation model, a reactive moderation model, an automated moderation model, a user-based moderation model and the like. For example, the pre-moderation model may be applied for content review and moderation before the content goes live. The post-moderation may be applied on a published multimedia content for review by the moderator I reviewer. The reactive moderation model may be applied on a published multimedia content with user feedback, where the moderator may review the content with the user feedback only. The automated moderation model may be an Al based model that may identify and filter the set of unsuitable segments of the multimedia content 110A automatically. The user-based moderation model may be configured to receive user inputs such as collaborative filtering, upvotes, downvotes, or content ranking mechanisms for the multimedia content 110A displayed to the user 114A. Further, the user-based moderation model may be configured to self-moderate the multimedia content 110A based on the user inputs.
[0025] In another embodiment, the content moderation model 102A may process a video stream. To process the video stream, the content moderation model 102A may use semantic segmentation to identify and potentially blur or mask inappropriate visual elements. A video optical character recognition (OCR) model associated with the content moderation model 102A may be used to detect and analyze text within the video, byflagging any problematic language. The content moderation model 102A may also include an audio content analyzer and an emotion classifier. The audio content analyzer and emotion classifier may evaluate the spoken content and emotional tone in the multimedia content 110A, to ensure that the audio aligns with the user's preferences and regional standards. In a family viewing scenario, the content moderation model 102A may dynamically adjust content based on the detected presence of children, automatically filtering out segments containing mature themes or language.
[0026] For instance, the content moderation model 102A may be adaptable. Thus, the adaptability allows the content moderation model 102A to handle diverse content types and viewing environments. For example, in an automotive setting, the content moderation model 102A may adjust content filtering based on the vehicle's location and driving conditions and prioritize audio content when the vehicle is in motion. In extended reality (XR) applications, the content moderation model 102A may apply content control to specific layers or elements within the XR space and provide a tailored and appropriate experience for the user 114A of different ages or sensitivities. The flexibility and contextawareness of the content moderation model 102A may make the content moderation model 102A a powerful tool to ensure safe, personalized, and pleasurable content experiences across various platforms and devices.
[0027] The set of sensors 102B may include suitable logic, circuitry, interfaces, and / or code that may be configured to capture the sensor data associated with the user 114A. The sensor data may include, but is not limited to, location data, proximity data, biometric data, image data, interactive data, audio data, or haptic data associated with user 114A. In an embodiment, the set of sensors 102B may include at least one of a camera, amicrophone, or a biometric sensor. The sensor data may be used to determine an age of the user 114A, which may be used for determination of the profile information associated with the user 114A.
[0028] In an embodiment, the set of sensors 102B may include the camera. The camera may detect the user 114A who operates the user device 114 associated with the first electronic device 102. The camera may capture an image of the user 114A. In another embodiment, the set of sensors 102B may include the microphone. The microphone may detect the audio associated with the user 114A. Further, the microphone may store the detected audio of the user 114A. In another embodiment, the set of sensors 102B may include the biometric sensor. The biometric sensor may detect the biometric data associated with the user 114A. For example, the biometric data may include, at least one of, but not limited to, an electroencephalogram (EEG), a magnetocardiogram (MCG), an electrocardiogram (ECG), or an electromyogram (EMG) associated with the user 114A. The biometric data may be used to determine the age, health condition, or the likes associated with the user 114A. Example implementations of the set of sensors 102B may include, but are not limited to, a proximity sensor, a motion detector, an optical detector, or a gesture recognition sensor.
[0029] For example, a camera from the set of sensors 102B may be used to detect the presence and estimate the age of viewers in the vicinity of the first electronic device 102. An audio sensor might be employed to analyze ambient sounds and determine if multiple users are present. A biometric sensor, such as an iris scanner or an ocular scanner may be used to determine more accurate age of viewers and enhance an ability of the content moderation model 102A to apply appropriate content filters.
[0030] The data collected by the set of sensors 102B may be utilized for various purposes for the content moderation. For instance, the first electronic device 102 may use visual content to dynamically adjust content filtering based on the detected presence of children. The audio content may be used to modulate volume levels or switch to more family-friendly content when multiple viewers are detected. The biometric content data may enable the first electronic device 102 to apply user-specific content preferences automatically and enhance the personalization of the viewing experience.
[0031] The second electronic device 104 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive the multimedia content 110A and the content metadata 110B associated with the plurality of segments of the multimedia content 110A from a content server, such as the server 106. In an embodiment, the content server may be associated with a VOD server, an Over-The-Top (OTT) server, a Television (TV) broadcasting server, and the like. The second electronic device 104 may include the VFM 104A, the NLM 104B, and the ALM 104C. The second electronic device 104 may be configured to apply the VFM 104A to analyze visual content of the multimedia content 110A. For example, the visual content of the multimedia content 110A may be image content, video content, animation content, art content, augmented reality (AR) content, virtual reality (VR) content, subtitle information, close caption content, and the like. The second electronic device 104 may determine the context information based on the analysis of the visual content and transmit the context information to the first electronic device 102. Examples of the second electronic device 104 may include, but are not limited to, a computing device, a server, a network provider, a base station, a router, a smartphone, a cellular phone, a mobile phone, a gaming device,a mainframe machine, a computer workstation, a consumer electronic (CE) device, a television, a set-top box, a broadcaster, a base station device, and / or the likes.
[0032] The second electronic device 104 may apply a language model (LM) to analyze audio content or text content of the multimedia content 110A to determine the context information. The audio content may be a background music, a sound effect, an audiobook, a podcast, an interactive audio, an audio track, or the likes. The text content may be subtitles, captions, scripts, articles, blogs, or the likes. The second electronic device 104 may transmit the context information to the first electronic device 102. The LM may correspond to at least one of the NLM 104B or the ALM 104C.
[0033] In an embodiment, the second electronic device 104 may be configured to segment the multimedia content 110A into the plurality of segments. The second electronic device 104 may generate the content metadata 110B for each segment of the plurality of segments and transmit the multimedia content 110A and the content metadata 110B to the first electronic device 102. In an embodiment, the second electronic device 104 may be configured to retrieve the region-specific profanity information from the database 108 and transmit the region-specific profanity information to the first electronic device 102. The region-specific profanity information may correspond to information associated with at least one of, but not limited to, offensive content criteria, inappropriate content criteria, or vulgar language that may be particular to a certain geographical region, culture, or linguistic community. For example, the unsuitable segment of the multimedia content 110 may be a harmless phrase in one region and may be deeply offensive in another region, due to cultural norms, traditions, or sensitivities.
[0034] The VFM 104A may include suitable logic, circuitry, interfaces, and / or code thatmay be configured to analyze the visual content of the multimedia content 110A to determine the context information associated with the visual content. For example, the VFM 104A may be a Al model that may process, understand, and generate visual data. The VFM 104A may be trained on extensive datasets of images and videos. The VFM 104A may perform various vision-related tasks such as an object detection, an image classification, segmentation of the multimedia content 110A, the content metadata generation, and image generation. The VFM 104A may be pretrained on diverse visual datasets and may be fine-tuned for specialized tasks through transfer learning. The transfer learning may be a machine learning (ML) technique where the VFM 104A model trained on one task or domain may be repurposed and fine-tuned to perform a different but related task. The VFM 104A may integrate with language models (LM) such as the NLM 104B or the ALM 104C for multimodal capabilities and enabling image captioning and answering questions about visual content.
[0035] For example, the VFM 104A may be used to identify potentially inappropriate or sensitive visual elements. The VFM 104A may detect and classify objects, recognize facial expressions, identify explicit content, and understand the overall context of visual scenes. For example, the VFM 104A may flag violent imagery, detect nudity, or identify age-inappropriate visual content in videos or images.
[0036] The NLM 104B may include suitable logic, circuitry, interfaces, and / or code that may be configured to analyze text content of the multimedia content 110A to determine context information associated with the multimedia content 110A. For example, the NLM 104B may be a Large Language Model (LLM) that may understand, generate, and process human language such as text content. Further, the NLM 104B may be trained ondatasets corresponding to the content metadata 11 OB associated with the multimedia content 110A, the NLM 104B may perform tasks like interpretation of text, generation of coherent content, translation of languages, answering questions, and engaging in conversations using a deep learning technique. The deep learning technique may include a neural networks and attention mechanisms to determine the context information and predict text sequences associated with the text content.
[0037] For example, the NLM 104B may play a crucial role in analysis of textual elements of the multimedia content 110A. the NLM 104B may be used to detect inappropriate language, identify hate speech, recognize potential threats, and understand the overall tone and context of written or transcribed content. The NLM 104B may be particularly useful in moderation of user-generated content, such as comments, subtitles, or chat messages in live streaming platforms. The NLM 104B may be applied in various scenarios. For instance, in a social media platform, the NLM 104B may analyze post captions and comments to flag potentially offensive or harmful content. In video streaming services, the NLM 104B may process subtitle tracks to ensure age-appropriate language. In online gaming environments, The NLM 104B may monitor in-game chat to prevent bullying or inappropriate communication.
[0038] The ALM 104C may include suitable logic, circuitry, interfaces, and / or code that may be configured to analyze audio content of the multimedia content 110A to determine context information associated with the multimedia content 110A. For example, the ALM 104C may be based on a neural network that may analyze and understand spoken language using a deep learning technique. The ALM 104C may be built with architecture like Recurrent neural networks (RNNs), convolutional neural network (CNNs), orTransformers. The ALM 104C may capture temporal patterns in audio data such as waveforms or spectrograms. Further, the ALM 104C may be trained on datasets corresponding to the content metadata 11 OB associated with the multimedia content 110A. The datasets may include a paired audio content and text content. Further, the ALM 104C may learn the relationship between spoken language of the audio content and written language of the text content. In an embodiment, the ALM 104C may include automatic speech recognition (ASR), text-to-speech (TTS), speaker identification, emotion recognition, and real-time audio-driven translation.
[0039] For example, the ALM 104C may be crucial for analysis of the audio components of multimedia content. The ALM 104C may be used to detect inappropriate language in spoken content, identify explicit or violent sound effects, recognize emotional distress in voices, and understand the overall audio context. The ALM 104C may be particularly valuable in moderation of content where visual cues are limited or absent, such as podcasts, voice messages, or background audio in videos. The ALM 104C may find application in various content moderation scenarios. In a video sharing platform, the ALM 104C may analyze the audio track to flag instances of profanity or hate speech that may not be captured in subtitles. In music streaming services, the ALM 104C may screen song lyrics for age-inappropriate content. In voice-based social media platforms, the ALM 104C may monitor live audio streams to ensure compliance with community guidelines.
[0040] In accordance with an embodiment, each of the VFM 104A, NLM 104B, and ALM 104C may be a neural network model having a plurality of layers with each layer forming a loop where the outputs of each element feed into the other elements, gradually. The loop of providing outputs of each element as inputs to other element for each of theVFM 104A, NLM 104B, and ALM 104C improve determination of the voice captions and the non-voice captions. The plurality of layers of the neural network model may include an input layer, one or more hidden layers, and an output layer. Each layer of the plurality of layers may include one or more nodes (or artificial neurons, represented by circles, for example). Outputs of all nodes in the input layer may be coupled to at least one node of hidden layer(s). Similarly, inputs of each hidden layer may be coupled to outputs of at least one node in other layers of the neural network model. Outputs of each hidden layer may be coupled to inputs of at least one node in other layers of the neural network model. Node(s) in the final layer may receive inputs from at least one hidden layer to output a result. The number of layers and the number of nodes in each layer may be determined from hyper-parameters of the neural network model. Such hyper-parameters may be set before, while training, or after training the neural network model on a training dataset.
[0041] Each node of the neural network model may correspond to a mathematical function (e.g. , a sigmoid function or a rectified linear unit) with a set of parameters, tunable during training of the neural network model. The set of parameters may include, for example, a weight parameter, a regularization parameter, and the like. Each node may use the mathematical function to compute an output based on one or more inputs from nodes in other layer(s) (e.g., previous layer(s)) of the neural network model. All or some of the nodes of the neural network model may correspond to same or a different mathematical function.
[0042] In training of the neural network model, one or more parameters of each node of the neural network model may be updated based on whether an output of the final layer for a given input (from the training dataset) matches a correct result based on a lossfunction for the neural network model. The above process may be repeated for the same or a different input until a minima of loss function may be achieved, and a training error may be minimized. Several methods for training are known in art, for example, gradient descent, stochastic gradient descent, batch gradient descent, gradient boost, metaheuristics, and the like.
[0043] The neural network model may include electronic data, which may be implemented as, for example, a software component of an application executable on the second electronic device 104. The neural network model may rely on libraries, external scripts, or other logic / instructions for execution by a processing device, such as, the second electronic device 104. The neural network model may include code and routines configured to enable a computing device to perform one or more operations. Additionally, or alternatively, the neural network model may be implemented using hardware including a processor, a microprocessor, a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in some embodiments, the neural network model may be implemented using a combination of hardware and software.
[0044] The server 106 may include suitable logic, circuitry, and interfaces, and / or code that may be configured to execute operations, such as data / file storage, rendering the multimedia content 110A, removal of the set of unsuitable segments, generation and playback of the moderated media content based on the multimedia content 110A, and the removal of the set of unsuitable segments. In one or more embodiments, the server 106 may store the multimedia content 110A and the content metadata 110B and may execute at least one operation associated with the first electronic device 102. The server 106 maybe implemented as a cloud server and may execute operations through web applications, cloud applications, HTTP requests, repository operations, file transfer, and the like. Other example implementations of the server 106 may include, but are not limited to, a database server, a file server, a web server, a media server, an application server, a mainframe server, or a cloud computing server.
[0045] In at least one embodiment, the server 106 may be implemented as a plurality of distributed cloud-based resources by use of several technologies that are well known to those ordinarily skilled in the art. A person with ordinary skill in the art will understand that the scope of the disclosure may not be limited to the implementation of the server 106, the first electronic device 102 and the second electronic device 104, as three separate entities. In certain embodiments, the functionalities of the server 106 may be incorporated in its entirety or at least partially in the first electronic device 102 or the second electronic device 104 without a departure from the scope of the disclosure. In certain embodiments, the server 106 may host the database 108. Alternatively, the server 106 may be separate from the database 108 and may be communicatively coupled to the database 108.
[0046] The database 108 may include suitable logic, interfaces, and / or code that may be configured to store the multimedia content 110A and the content metadata 110B (and / or the moderated media content). The database 108 may also store region-specific profanity information associated with a geographic location of the first electronic device 102 and / or the second electronic device 104. The database 108 may be derived from data of a relational or non-relational database or a set of comma-separated values (csv) files in conventional or big-data storage. The database 108 may be stored or cached ona device, such as a server (e.g., the server 106), the first electronic device 102, the second electronic device 104. The device storing the database 108 may be configured to receive a query for the multimedia content 110A and the content metadata 110B from the first electronic device 102 or the second electronic device 104. Based on the received query, the device that stores the database 108 may retrieve and provide the multimedia content 110A and the content metadata 110B to the first electronic device 102, or the second electronic device 104.
[0047] In some embodiments, the database 108 may be hosted on a plurality of servers stored at the same or different locations. The operations of the database 108 may be executed using hardware, including a processor, a microprocessor (e.g., to perform or control performance of one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). In some other instances, the database 108 may be implemented using software. In some embodiment, the functionalities of the database 108 may be implemented by the server 106 and / or the first electronic device 102, without departure from the scope of the disclosure.
[0048] The multimedia content 110A may include at least one of broadcast content, video-on-demand content, or streaming content. In an embodiment, the multimedia content 110A may be image content, video content, animation content, an art, augmented reality (AR) content, virtual reality (VR) content, subtitle information, close caption content, and the likes. The content metadata 110B may include at least one of a title, a description, a format, dimensions, a color space, a date, a location, an author / photographer, a license, an attribution resolution, a frame rate, a director / producer, a language, a genre, keywords, an album, a duration, a file size, a category, and the likes,associated with the multimedia content 110A.
[0049] The communication network 112 may include a communication medium through which the first electronic device 102, the second electronic device 104, the server 106, and the user device 114 may communicate with one another. The communication network 112 may be one of a wired connection or a wireless connection. Examples of the communication network 112 may include, but are not limited to, the Internet, a cloud network, Cellular or Wireless Mobile Network (such as Long-Term Evolution and 5th Generation (5G) New Radio (NR)), a Wireless Fidelity (Wi-Fi) network, a satellite network (e.g., using a network of low earth orbit satellites), a Personal Area Network (PAN), a Local Area Network (LAN), or a Metropolitan Area Network (MAN). Various devices in the network environment 100 may be configured to connect to the communication network 112 in accordance with various wired and wireless communication protocols. Examples of such wired and wireless communication protocols may include, but are not limited to, at least one of a Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Zig Bee, EDGE, IEEE 802.11 , light fidelity (Li-Fi), 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device to device communication, cellular communication protocols, and Bluetooth (BT) communication protocols.
[0050] The user device 114 may be associated with the first electronic device 102, or the second electronic device 104 and may include suitable logic, circuitry, interfaces, and / or code that may be configured to render the multimedia content 110A and the content metadata 110B along with the moderated media content. The first electronicdevice 102 or the second electronic device 104 may control the user device 114 to playback or render the multimedia content 110A and the content metadata 110B and the generated moderated media content. In certain embodiments, the user device 114 may upload (for example, based on a user-input) the multimedia content 110A and the content metadata 11 OB to the database 108 for storage. Additionally, or alternatively, the user device 114 may transmit the multimedia content 110A and the content metadata 110B to the first electronic device 102, or the second electronic device 104.
[0051] In an embodiment, the user device 114 may be controlled to render the moderated media content and playback the moderated media content on a display device associated with the user device 114, or the first electronic device 102. In an embodiment, the user device 114 may include a user interface that may allow a user 114A to interact with the user device 114. The user device 114 may receive a user input, through the user interface, to modify and update the content moderation model 102A. The user device 114 may transmit such user input to the first electronic device 102, or the second electronic device 104. In embodiment, the user 114A may be a person in the proximity of the user device 114 or a person who operates the user device 114 to watch the multimedia content 110A.
[0052] In operation, the first electronic device 102 may be configured to receive the multimedia content 110A and the content metadata 110B associated with a plurality of segments of the multimedia content 110A from the second electronic device 104. The plurality of segments of the multimedia content 110A may include diverse components or parts that may be combined to generate the multimedia content 110A. For example, the plurality of segments may be a text segment, a visual segment, audio segment, a videosegment, a metadata segment, or the likes.
[0053] The first electronic device 102 may be configured to determine context information associated with each segment of the plurality of segments, based on the content metadata. The context information may be supplementary details that provide insights into the multimedia content characteristics, purpose, or background. The context information may be used for interpretation and analysis of multimedia content 110A based on additional data such as content metadata 110B.
[0054] The first electronic device 102 may be configured to receive the region-specific profanity information associated with the multimedia content 110A from the second electronic device 104. In an embodiment, the first electronic device 102 may be configured to determine the geographic location of the first electronic device 102 and receive the region-specific profanity information based on the determined geographic location. The region-specific profanity information may correspond to information associated with at least one of, but not limited to, offensive content criteria, inappropriate content criteria, or vulgar language that may be particular to a certain geographical region, culture, or linguistic community. In an embodiment, the region-specific profanity information associated with the multimedia content 110A may be received from the server 106 by the first electronic device 102 or the second electronic device 104.
[0055] The first electronic device 102 may be configured to receive, from the set of sensors 102B, the sensor data associated with the user 114A present in the predetermined proximity of the first electronic device. For example, the predetermined proximity may be a set distance or a range within which the set of sensors 102B may detect an object or a user 114A, and measure a change, or respond to a stimulus. Thepredetermined proximity may be defined based on a specification of the set of sensors 102B and an intended application of the set of sensors 102B. In an embodiment, the first electronic device 102 may be configured to determine an age of the user 114A based on the sensor data.
[0056] The first electronic device 102 may be configured to determine profile information associated with the user 114A based on the sensor data. The set of sensors 102B may Include at least one of: a camera, a microphone, or a biometric sensor. For example, the profile information may include, but is not limited to, personal details, preferences, and attributes associated with of the user 114A. The profile information may describe the user 114A associated within the first electronic device 102 or the user device 114. Further, the personal details may include name, age, gender, contact information, date of birth, location, and the likes. Additionally, the profile information may also include demographic and behavioral insights generated based on aggregated data and optional user inputs like surveys. In an embodiment, the first electronic device 102 may be configured to determine profile information associated with the user 114A based on the age of the user 114A. In an embodiment, the first electronic device 102 may be configured to detect the change in the sensor data to update the profile information based on the change in the sensor data.
[0057] The first electronic device 102 may be configured to apply the content moderation model 102A on the plurality of segments, based on the context information, the region-specific profanity information, and the profile information. Further, the first electronic device 102 may be configured to receive a user input to modify the content moderation model 102A and update the content moderation model 102A based on theuser input.
[0058] The first electronic device 102 may be configured to select, from the plurality of segments, the set of unsuitable segments for the user, based on the application of the content moderation model 102A. The first electronic device 102 may be configured to mask or replace the set of unsuitable segments in the moderated media content.
[0059] For example, the VFM 104A may employ a multi-stage processing pipeline on the multimedia content 110A and the content metadata 110B in the content moderation model 102A. The VFM 104A may use convolutional neural networks to extract low-level features from images or video frames. The features may then be processed through higher-level networks to recognize objects, scenes, and actions. The VFM 104A may also incorporate semantic segmentation techniques to understand the spatial relationships between different elements in a scene. The results of this analysis may be combined with content metadata 110B and context information to make informed decisions about content suitability.
[0060] For example, the operation of the NLM 104B in the content moderation model 102A may involve several processes. The NLM 104B may tokenize the input text into manageable units. Using techniques like transformer architectures or recurrent neural networks, the NLM 104B may process the tokens to extract the context and meaning. The NLM 104B may also incorporate sentiment analysis to extract the emotional tone of the text. Additionally, the NLM 104B may use named entity recognition to identify specific references that might require moderation. The NLM 104B may then combine results of the analysis to generate decisions about content suitability, potentially in conjunction with region-specific profanity databases and user profile information.
[0061] For example, the operation of the ALM 104C in the content moderation model 102A may involve multiple processes. The ALM 104C may convert the audio input into a spectrogram or other frequency-based representation. This representation may then be processed through convolutional or recurrent neural networks to extract relevant features. The ALM 104C may employ speech recognition techniques to transcribe spoken content, which can then be analyzed for inappropriate language. Simultaneously, the ALM 104C may use sound event detection to identify non-speech audio elements that may require moderation. The ALM 104C may also incorporate emotion recognition to understand the tone and sentiment of spoken content. The results from the various analyses may be combined to make comprehensive decisions about the suitability of audio content, potentially in integration with region-specific guidelines and user preferences.
[0062] The first electronic device 102 may be configured to determine the moderated media content based on the multimedia content 110A and the set of unsuitable segments. Further, the first electronic device 102 may be configured to dynamically update the moderated media content based on the updated profile information. In another embodiment, the first electronic device 102 may be configured to generate the content moderation report based on the set of unsuitable segments and transmit the content moderation report to the second electronic device 104, or the user device 114.
[0063] The first electronic device 102 may be configured to control the user device 114 to render the moderated media content. In an embodiment, the first electronic device 102 may be configured to generate a content recommendation based on the profile information and control the user device 114 to render the content recommendation.
[0064] Existing content moderation systems often face challenges in provision of contextually appropriate content based on personalized content filtering. Current solutions typically rely on pre-defined rating systems or content tags, which may not account for regional variations in language and cultural norms. These systems often lack the ability to adapt to individual user preferences or real-time changes in viewing environments. As a result, conventional approaches may lead to over-blocking of harmless content or under-filtering of inappropriate material, which may result in suboptimal user experiences and potential exposure to unsuitable content.
[0065] The disclosed content moderation technique utilizes a combination of advanced technologies to provide a more nuanced and adaptive approach to content filtering. By leveraging context information, region-specific profanity data, and user profile information, the disclosed technique may make more informed decisions about content suitability. Unlike traditional systems, the disclosed approach may dynamically adjust moderation parameters based on the specific characteristics of the content, regional norms, and individual user profiles. The disclosed technique may offer improved accuracy in identification of unsuitable content, better adaptation to regional and cultural sensitivities, and enhanced personalization of content moderation based on user profiles. Additionally, the disclosed technique may have an ability to dynamically update the moderation process based on real-time sensor data, which may allow a more responsive and context-aware content filtering experience.
[0066] FIG. 2A is a block diagram that illustrates an exemplary electronic device of FIG.1 , in accordance with an embodiment of the disclosure. FIG. 2A is explained in conjunction with elements from FIG. 1 . With reference to FIG. 2A, there is shown the firstelectronic device 102. The first electronic device 102 may include circuitry 202A, a memory 204A, an input / output (I / O) device 206A, a network interface 208A, the content moderation model 102A, and the set of sensors 102B. The input / output (I / O) device 206A may include a display device 210A. The memory 204A may further store the multimedia content 110A and the content metadata 110B
[0067] The circuitry 202A may include suitable logic, circuitry, and / or interfaces that may be configured to execute program instructions associated with different operations to be executed by the first electronic device 102. For example, the operations may include multimedia content and the content metadata reception, context information determination, region-specific profanity information and sensor data reception, profile information determination, the content moderation model 102A application, the set of unsuitable segments selection, the moderated media content determination, and control the user device 114 to render the moderated media content. The circuitry 202A may include one or more processing units, which may be implemented as a separate processor. In an embodiment, the one or more processing units may be implemented as an integrated processor or a cluster of processors that perform the functions of the one or more specialized processing units, collectively.
[0068] The circuitry 202A may be implemented based on a number of processor technologies known in the art. Examples of implementations of the circuitry 202A may be an X86-based processor, a Graphics Processing Unit (GPU), a Reduced Instruction Set Computing (RISC) processor, an Application-Specific Integrated Circuit (ASIC) processor, a Complex Instruction Set Computing (CISC) processor, a microcontroller, a central processing unit (CPU), and / or other control circuits.
[0069] The memory 204A may include suitable logic, circuitry, interfaces, and / or code that may be configured to store one or more instructions to be executed by the circuitry 202A (and / or the first electronic device 102) to perform the operations of the circuitry 202A (and / or the first electronic device 102). The memory 204A may be configured to store the multimedia content 110A, the content metadata 11 OB, the determined moderated media content, and data associated with the user device 114. Examples of implementation of the memory 204A may include, but are not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Hard Disk Drive (HDD), a Solid-State Drive (SSD), a CPU cache, and / or a Secure Digital (SD) card.
[0070] The I / O device 206A may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive an input and provide an output based on the received input. For example, the I / O device 206A may receive the multimedia content 110A and the content metadata 110B. The I / O device 206A may receive a user input to initiate content moderation. Further, the I / O device 206A may control the user device 114 to render the multimedia content 110A and the content metadata 110B and the determined moderated media content. The I / O device 206A may include the display device 210A. Examples of the I / O device 206A may include, but are not limited to, a touch screen, a keyboard, a mouse, a joystick, a microphone, or a speaker.
[0071] The network interface 208A may include suitable logic, circuitry, interfaces, and / or code that may be configured to facilitate communication between the first electronic device 102 and the server 106 via the communication network 112. The network interface 208A may be implemented by use of various known technologies tosupport wired or wireless communication of the first electronic device 102 with the communication network. The network interface 208A may include, but is not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, or a local buffer circuitry.
[0072] The network interface 208A may be configured to communicate via wireless communication with networks, such as the Internet, an Intranet, a wireless network, a cellular telephone network, a wireless local area network (LAN), or a metropolitan area network (MAN). The wireless communication may be configured to use one or more of a plurality of communication standards, protocols and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), wideband code division multiple access (W-CDMA), Long Term Evolution (LTE), 5thGeneration (5G) New Radio (NR), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (such as IEEE 802.11a, IEEE 802.11 b, IEEE 802.11g or IEEE 802.11 n), voice over Internet Protocol (VoIP), light fidelity (Li-Fi), Worldwide Interoperability for Microwave Access (Wi-MAX), a protocol for email, instant messaging, and a Short Message Service (SMS).
[0073] The display device 210A may include suitable logic, circuitry, and interfaces that may be configured to display the moderated media content and the multimedia content 110A and the content metadata 110B. The display device 210A may be a touch screen which may enable the user 114A to provide a user-input via the display device 210A. The touch screen may be at least one of a resistive touch screen, a capacitive touch screen, or a thermal touch screen. The display device 210A may be realized through severalknown technologies such as, but not limited to, at least one of a Liquid Crystal Display (LCD) display, a Light Emitting Diode (LED) display, a plasma display, or an Organic LED (OLED) display technology, or other display devices. In accordance with an embodiment, the display device 21 OA may refer to a display screen of a head mounted device (HMD), a smart-glass device, a see-through display, a projection-based display, an electro- chromic display, or a transparent display.
[0074] FIG. 2B is a block diagram that illustrates an exemplary electronic device of FIG.1 , in accordance with an embodiment of the disclosure. FIG. 2B is explained in conjunction with elements from FIG. 1. With reference to FIG. 2B, there is shown the Second electronic device 104. The Second electronic device 104 may include circuitry 202B, a memory 204B, an input / output (I / O) device 206B, a network interface 208B, the VFM 104A, the NLM 104B, and the ALM 104C. The input / output (I / O) device 206B may include a display device 21 OB.
[0075] The circuitry 202B may include suitable logic, circuitry, and / or interfaces that may be configured to execute program instructions associated with different operations to be executed by the second electronic device 104. For example, the operations may include multimedia content reception, the VFM application, the NLM application, the ALM application, the multimedia content segmentation, content metadata generation, and the region-specific profanity information retrieval. The circuitry 202B may include one or more processing units, which may be implemented as a separate processor. In an embodiment, the one or more processing units may be implemented as an integrated processor or a cluster of processors that perform the functions of the one or more specialized processing units, collectively.
[0076] The circuitry 202B may be implemented based on a number of processor technologies known in the art. Examples of implementations of the circuitry 202B may be an X86-based processor, a Graphics Processing Unit (GPU), a Reduced Instruction Set Computing (RISC) processor, an Application-Specific Integrated Circuit (ASIC) processor, a Complex Instruction Set Computing (CISC) processor, a microcontroller, a central processing unit (CPU), and / or other control circuits.
[0077] The memory 204B may include suitable logic, circuitry, interfaces, and / or code that may be configured to store one or more instructions to be executed by the circuitry 202B (and / or the second electronic device 104) to perform the operations of the circuitry 202B (and / or the second electronic device 104). The memory 204B may be configured to store the moderated media content, and the multimedia content 110A, the content metadata 110B and data associated with the user device 114. Examples of implementation of the memory 204B may include, but are not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Electrically Erasable Programmable Read- Only Memory (EEPROM), Hard Disk Drive (HDD), a Solid-State Drive (SSD), a CPU cache, and / or a Secure Digital (SD) card.
[0078] The I / O device 206B may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive an input and provide an output based on the received input. For example, the I / O device 206B may receive the multimedia content 110A and the content metadata 110B. The I / O device 206B may control the user device 114 to render the multimedia content 110A and the content metadata 110B and the determination of the moderated media content. The I / O device 206B may include the display device 210B. Examples of the I / O device 206B may include, but are not limitedto, a touch screen, a keyboard, a mouse, a joystick, a microphone, or a speaker.
[0079] The network interface 208B may include suitable logic, circuitry, interfaces, and / or code that may be configured to facilitate communication between the second electronic device 104 and the server 106 via the communication network 112. The network interface 208B may be implemented by use of various known technologies to support wired or wireless communication of the Second electronic device 104 with the communication network. The network interface 208B may include, but is not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, or a local buffer circuitry.
[0080] The network interface 208B may be configured to communicate via wireless communication with networks, such as the Internet, an Intranet, a wireless network, a cellular telephone network, a wireless local area network (LAN), or a metropolitan area network (MAN). The wireless communication may be configured to use one or more of a plurality of communication standards, protocols and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), wideband code division multiple access (W-CDMA), Long Term Evolution (LTE), 5thGeneration (5G) New Radio (NR), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (such as IEEE 802.11a, IEEE 802.11 b, IEEE 802.11g or IEEE 802.11 n), voice over Internet Protocol (VoIP), light fidelity (Li-Fi), Worldwide Interoperability for Microwave Access (Wi-MAX), a protocol for email, instant messaging, and a Short Message Service (SMS).
[0081] The display device 210B may include suitable logic, circuitry, and interfaces thatmay be configured to display the moderated media content, and the multimedia content 110A and the content metadata 11 OB. The display device 21 OB may be a touch screen which may enable a user to provide a user-input via the display device 21 OB. The touch screen may be at least one of a resistive touch screen, a capacitive touch screen, or a thermal touch screen. The display device 21 OB may be realized through several known technologies such as, but not limited to, at least one of a Liquid Crystal Display (LCD) display, a Light Emitting Diode (LED) display, a plasma display, or an Organic LED (OLED) display technology, or other display devices. In accordance with an embodiment, the display device 21 OB may refer to a display screen of a head mounted device (HMD), a smart-glass device, a see-through display, a projection-based display, an electro- chromic display, or a transparent display.
[0082] FIG. 3 is a diagram that illustrates an exemplary scenario of a content moderation system, in accordance with at least one embodiment of the disclosure. FIG. 3 is described in conjunction with elements from FIG. 1 , FIG. 2A, and FIG. 2B. With reference to FIG. 3, there is shown an exemplary scenario 300 of a content moderation system. The content moderation system of the scenario 300 includes the first electronic device 102, the second electronic device 104, the user device 114, and the communication network 112. The user device 114 includes the first electronic device 102 a screening room 302, an audience software development kit (SDK) 304, and the set of sensors 102B. The screening room 302 includes a content playback module 302A. The audience SDK 304 includes a sensor module 304A and a detection module 304B. In an embodiment, a circuitry associated with the user device 114 may be configured to host the screening room 302, and the audience SDK 304. In another embodiment, the circuitry202B of the second electronic device 104 may be configured to host the screening room302, and the audience SDK 304.
[0083] In the scenario 300, the first electronic device 102 may receive the multimedia content 110 and the content metadata 110B associated with the multimedia content 110A from the second electronic device 104 via the communication network 112. The content metadata 110B may include information such as scene descriptions, dialogue transcripts, audio characteristics, and visual element classifications. The content playback module 302A within the screening room 302 may manage the playback of the received multimedia content 110A and coordinate with other components to ensure appropriate content moderation and delivery.
[0084] The screening room 302 may be a foreground application of the user device 114 that may be configured to create a theater-like viewing experience for the user 114A. Further, the circuitry 202B may configure the screening room 302 to provide features that enhance the playback of the multimedia content 110A such as watching movies, TV shows, or other multimedia content in a shared or immersive environment. The enhanced features may be streaming high-quality content, offering early access to movie releases, which may facilitate virtual watch parties, and the likes.
[0085] The content playback module 302A of the screening room 302 may be a framework to stream, deliver, and consume multimedia content. The content playback module 302A may offer a seamless and immersive viewing experience. The features offered by the content playback module 302A may include on-demand and interactive playback, which may allow the user 114A to control viewing with options to pause, rewind, and access subtitles or bonus content. The content playback module 302A may includeadaptive streaming optimizes playback quality based on network conditions using protocols like Dynamic Adaptive Streaming over HTTP (DASH) or HTTP Live Streaming (HLS). The content playback module 302A may incorporate Digital Rights Management (DRM) to ensure secure content delivery to the user 114A. The content playback module 302A supports multi-device accessibility, which may enable the user 114A to switch devices without interruption and enhance viewing quality with high-resolution video and advanced audio formats for a theater-like experience.
[0086] The audience SDK 304 may be a background application of the user device 114. The audience SDK 304 may be a collection of tools, libraries, and applications that may be provided by the user device 114. The circuitry 202B may configure the audience SDK 304 to integrate audience-related features into the user device 114. Further, the circuitry 202B may configure the audience SDK 304 to facilitate audience management based on identification and categorization of user groups based on demographics, behavior, or preferences, to enable targeted campaigns and personalized experiences. Further, the circuitry 202B may configure the audience SDK 304 to analyze the multimedia content 110A, or the profile information to extract insights associated with user engagement and behavior, to track key metrics like retention rates and session durations. The audience SDK 304 may include user engagement tools such as push notifications and in-app messages to enhance interaction and loyalty. Integration capabilities of the audience SDK 304 may ensure compatibility with third-party tools and various platforms, while privacy and compliance features may ensure data security and adherence to regulations. For example, the audience SDK 304 may widely be used in marketing,media, gaming, and entertainment to optimize user experiences and create data-driven interactions.
[0087] In an embodiment, the audience SDK 304 may include a detection module 304B, and a sensor module 304A. The detection module 304B may be communicably coupled with the set of sensors 102B such as a camera. The sensor module 304A of the audience SDK 304 may be configured to capture an image or a video of the user 114A based on the detection of the user 114A in the proximity of the user device 114. The sensor module 304A of the audience SDK 304 may be configured to use the sensor data to detect, analyze, and interpret audience-related data through the ML technique or the VFM 104A. The sensor module 304A may identify an individual or groups of individuals in the camera’s view, and detect faces, bodies, or gestures. The sensor module 304A may analyze attributes like age, gender, facial expressions, and engagement levels, to generate insights into audience interaction with content or environments. The sensor module 304A may further be configured to track behaviors such as gaze direction and movement patterns to measure attention or interest.
[0088] The detection module 304B of the audience SDK 304 may be configured to identify, classify, and segment audience I user data based on a predefined parameters or patterns. The detection module 304B may collects data from various sources, such as the set of sensors 102B. for example, the data collected may be user activities, preferences, and interactions, including device information and browsing behavior. In an embodiment, the detection module 304B may use the ML techniques to recognize trends, clusters, and outliers in the data. The detection module 304B categorizes audiences into specific groups based on demographics, interests, or behaviors, to enable personalizedcontent delivery and targeted messaging. The detection module 304B may be configured to detect the presence of the user 114A to trigger immediate actions and may further use predictive modeling to forecast future behaviors and optimize audience engagement strategies. The detection module 304B may essentially be used to enhance user experiences, driving engagement, and support data-driven decision-making.
[0089] The first electronic device 102 may determine the context information for each segment of the multimedia content 110 based on the content metadata 11 OB. The process of determination of the context information may include the application of the content moderation model 102A, that may leverage the capabilities of the VFM 104A, the NLM 104B, and the ALM 104C for comprehensive analysis. The VFM 104A may identify visual elements such as objects, actions, and environments, while the ALM 104C may analyze audio characteristics including tone, emotion, and speech content. The NLM 104B may process textual data to understand themes, sentiment, and contextual nuances. Based on a combination of the analyses, the first electronic device 102 may derive rich context information, such as the emotional tone of a scene, the presence of potentially sensitive topics, or the overall narrative arc of the content.
[0090] For example, the audience SDK 304 may work in tandem with the set of sensors 102B to gather viewer information. The sensor module 304A may collect various types of raw data, which may include visual imagery from cameras, audio input from microphones, and biometric data from specialized sensors. The detection module 304B may process the collected raw data to determine viewer characteristics. For instance, a facial recognition algorithm may be used to estimate a viewer's age, an emotion recognition technique may be used to gauge the viewer’s reactions to content, and a voice analysistechnique may be used to detect multiple viewers in the room. Based on the sensor data, the first electronic device 102 may construct a detailed profile information. The profile information may include attributes such as age, viewing preferences, emotional responses to different content types, and potentially even cultural background inferred from language or accent detection. The content moderation model 102A of the first electronic device 102 may use the profile information to customize the moderation process for each specific viewer or group of viewers.
[0091] The application of the content moderation model 102A on the plurality of segments may be based on context information, the region-specific profanity information, and the profile information. A multi-faceted approach may allow for nuanced content filtering. For example, if the profile information of the user 114A indicates that the user 114A is a teenager, the first electronic device 102 may allow mild profanity but flag segments with extreme violence. For younger children, even mild thematic elements may be identified for potential moderation.
[0092] In an exemplary embodiment, the first electronic device 102 may employ various techniques to handle unsuitable segments. In some cases, the visual content may be blurred or pixelated to obscure graphic elements while narrative continuity may be maintained. In other instances, alternative scenes or footage may be seamlessly substituted. For the audio content, the first electronic device 102 may use sound mixing techniques to mute or replace inappropriate language while the overall audio quality may be preserved.
[0093] In an exemplary embodiment, the first electronic device 102 may dynamically generated accessibility features, based on viewer needs. For example, in case the user114A has hearing impairments, the first electronic device 102 may not only generate enhanced subtitles but also visual representations of important audio cues, such as sound effect icons or color-coded speaker identification. In case the user 114A is visually impaired, the first electronic device 102 may generate detailed audio descriptions of visual scenes, using natural language processing to create coherent narratives that fit within dialogue gaps.
[0094] In extended reality (XR) environments, the content moderation may become more complex. For non-head-mounted display (HMD) based XR systems, the first electronic device 102 may apply content control to specific layers or elements within the XR space. The first electronic device 102 may allow for selective moderation of interactive elements, background environments, or virtual characters while preserving the overall XR experience.
[0095] The first electronic device 102 adaptation to automotive scenarios may involve real-time adjustments based on driving conditions and passenger profiles. When a vehicle is in motion, the first electronic device 102 may prioritize audio content and limit visual distractions. Location-based content filtering may be applied, such as restricting news about traffic accidents when the vehicle approaches high-risk areas. For family vehicles, the first electronic device 102 may create separate audio zones, delivering age- appropriate content to different passengers simultaneously.
[0096] To enhance the moderation process, the content moderation model 102A may dynamically augment the content metadata 110B. The dynamically augment may include continuous analysis of the multimedia content 110A using the VFM 104A, NLM 104B, and ALM 104C. As the streaming and play back of the multimedia content 110A progresses,the content moderation model 102A of the first electronic device 102 may update analysis of themes, emotional arcs, and potential sensitivities. The dynamic approach may allow the first electronic device 102 to adapt to evolving narratives and ensure that multimedia content 110A may be evaluated in the content moderation model 102A, based on full context rather than relying solely on pre-defined tags or ratings.
[0097] The incorporation of spatial audio techniques may further refine the content moderation capabilities. The first electronic device 102 may manipulate the audio content of the multimedia content 110A to create directional sound, effectively "steering" certain audio elements away from specific listeners. The manipulation may particularly be useful in shared viewing environments, where adults and children may be watching together. Additionally, the content moderation model 102A may employ frequency masking or selective audio cancellation to obscure unsuitable audio content without disrupting the overall soundscape.
[0098] Based on integration of these advanced components and techniques, the content moderation model 102A may offer a highly adaptive and context-aware approach to content filtering. The content moderation model 102A may continuously analyze and respond to the interplay between content characteristics, viewer profiles, and environmental factors. This may enable the delivery of appropriately moderated experiences across a wide range of devices and scenarios, from traditional home entertainment systems to cutting-edge XR environments and smart vehicles.
[0099] FIG. 4 is a diagram that illustrates an exemplary scenario of a content moderation system, in accordance with an embodiment of the disclosure. FIG. 4 is described in conjunction with elements from FIG. 1 , FIG. 2A, FIG. 2B, and FIG. 3. Withreference to FIG. 4, there is shown an exemplary scenario 400 of a content moderation system. The scenario 400 includes the first electronic device 102, the second electronic device 104, and the user device 114. The first electronic device 102 includes the content moderation model 102A including a context controller 402, a moderation workflow generator 404, a content recommendation engine 406, and a region-specific moderation engine 408. The second electronic device 104 includes the VFM 104A including operations such as, a semantic segmentation 410, a spatial audio context 412, a content filtering 414, a content masking 416, and a video OCR 418. The second electronic device 104 may further include the NLM 104B and the ALM 104C, which may be associated with an audio content analyzer 420, an emotion classifier 422, and a search recommender 424.
[0100] In the scenario 400, the first electronic device 102 may receive multimedia content 110A and content metadata 110B from the second electronic device 104 via the communication network 112. The context controller 402 may be configured to analyze the received content metadata 110B and determine the context information for each segment of the multimedia content 110A. The context information may include factors such as scene type, dialogue content, and audio characteristics.
[0101] In an embodiment, the context controller 402 of the content moderation model 102A associated with the first electronic device 102 may be configured to analyze and interpret the context information of the multimedia content 110A. The context controller 402 may use contextual cues such as language (to assess cultural or situational nuances), behavior (to identify patterns like spamming or hate speech), and media (todistinguish between harmful and appropriate imagery) to determine whether content violates rules or guidelines.
[0102] The moderation workflow generator 404 of the content moderation model 102A associated with the first electronic device 102 may be configured to create a customized content moderation plan based on the context information and profile information associated with the user 114A. For example, if the profile information indicates a young viewer I user, the moderation workflow generator 404 may set stricter filtering parameters for violence and mature themes.
[0103] The set of sensors 102B of the first electronic device 102 may continuously monitor the viewing environment. In some cases, the first electronic device 102 may detect a change in the sensor data. For instance, the set of sensors 102B may detect the presence of a new viewer I user 114A or a change in the existing viewer's emotional state. Based on the change detected, the first electronic device 102 may update the profile information. The content moderation model 102A may then use the updated profile information to dynamically adjust the content moderation in real-time.
[0104] The content recommendation engine 406 of the content moderation model 102A associated with the first electronic device 102 may leverage the user profile information and the context information to generate personalized content recommendations. For example, if the profile information of the user 114A indicates a preference for educational content, the content recommendation engine 406 may suggest the multimedia content 110A associated with documentaries or educational programs. The first electronic device 102 may control the user device 114 to render the content recommendation and present the content recommendation through the user device 114.
[0105] The region-specific moderation engine 408 of the content moderation model 102A associated with the first electronic device 102 may be configured assess and regulate the multimedia content 110A based on the cultural, legal, and linguistic nuances of a particular region. The region-specific moderation engine 408 may be configured to apply location-based multimedia content filtering rules. For example, in regions with stricter multimedia content regulations, the region-specific moderation engine 408 may apply more conservative moderation policies. Further, the first electronic device 102 may control the user device 114 to render the moderated media content that may be more conservatively moderated. For example, the region-specific moderation engine 408 may align the displayed plurality of segments of the multimedia content 110A with local norms, laws, and sensitivities, taking into account regional contexts.
[0106] Based on the analysis from the context controller 402, the moderation workflow generator 404, the content recommendation engine 406, and the region-specific moderation engine 408, the first electronic device 102 may generate the content moderation report. The content moderation report may detail the identified set of unsuitable segments and the applied content moderation model 102A. In some cases, the first electronic device 102 may transmit the content moderation report to the second electronic device 104, potentially for further analysis or refinement of the content analysis algorithms.
[0107] For example, the first electronic device 102 may include a user interface workflow to moderate content consumption by children and young adults. The user interface may allow parents or guardians to set viewing preferences, time limits, orcontent restrictions. The user interface of the first electronic device 102 may display realtime updates on content being viewed and any applied moderation actions.
[0108] In scenarios where the plurality of segments of the current multimedia content 110A may be deemed unsuitable, the first electronic device 102 may suggest alternative or age-appropriate content based on the profile information and the user preference based on the profile information. For example, if a violent scene is detected in a movie being watched by a young user, the first electronic device 102 may suggest a more familyfriendly alternative with similar themes or characters.
[0109] The second electronic device 104 may employ various specialized modules for content analysis. The second electronic device 104 may determine context information of the visual content associated with the multimedia content 110A based on the application of VFM 104A. The VFM 104A may include at least the semantic segmentation 410, the spatial audio context 412, the content filtering 414, the content masking 416, and the video optical character recognition (OCR) 418.
[0110] The semantic segmentation 410 may be include analysis of the visual content based on a segmentation of the visual content into meaningful regions or objects, by assignment of a specific category label to each segment of the image or the video. For example, the semantic segmentation 410 may enable detailed analysis of the image's structure, video structure, allowing the VFM 104A to differentiate between various objects, surfaces, or elements within a scene, which may be essential for tasks like object detection, scene recognition, and image interpretation. In an embodiment, the semantic segmentation 410 may be configured to identify and categorize different elements within video frames, which may allow more precise content filtering.
[0111] The spatial audio context 412 may be configured to analyze the audio landscape of the content, which may enable selective audio moderation. In an embodiment, the spatial audio context 412 may be configured to integrate spatial audio cues with visual content and enhance an ability of the VFM 104A to interpret environments in a three- dimensional space. Based on an analysis of a direction, distance, and characteristics of sounds in relation to visual elements, the spatial audio context 412 may enable more immersive and context-aware content analysis, which may particularly be useful for applications like augmented reality, robotics, and scene comprehension.
[0112] The content filtering 414 and the content masking 416 may be configured to modify the set of unsuitable segments from the plurality of segments. The content filtering 414 of the VFM 104A may identify, categorize, and / or remove visual content based on predefined rules or guidelines. The content filtering 414 may further include analysis of the multimedia content 110A such as an image content or a video content to detect specific attributes, objects, or patterns that may be considered harmful, inappropriate, or irrelevant. Thus, the VFM 104A may ensure that the moderated media content aligns with requirements of the user 114A. For instance, the content filtering 414 may identify segments including excessive violence.
[0113] The content masking 416 of the VFM 104A may be obscure, mask, replace or conceal the set of unstable segments from the plurality of segments specific areas based on the application of the content moderation model 102A. For example, the content masking 416 may mask a specific segment within an image or a video based on certain criteria associated with the region-specific profanity information and the profile information associated with the user 114A. The content masking 416 may be used to hide sensitive,private, or irrelevant information while the rest of the visual content may be retained for analysis or presentation. Thus, the content masking 416 may ensure privacy, compliance with ethical guidelines, and focused processing for tasks like anonymization or selective data handling. For instance, the content masking 416 may apply visual blurring or audio muting to the set of unstable segments.
[0114] The video OCR 418 of the VFM 104A may be configured to extract and analyze text content of the multimedia content 110A. The video OCR 418 may identify text content such as signs, labels, or subtitles, within the visual content (such as a video of the multimedia content 110A) and convert the text content of the visual content into machine- readable text content. For example, the video OCR 418 may extract and analyze text that appears in the video content and may enable the moderation of on-screen text or subtitles. Thus, the video OCR 418 may enable automated transcription, data indexing, or real-time information retrieval from the visual content or the video content.
[0115] The language model such as the NLM 104B or the ALM 104C may be configured to analyze audio content or the text content of the multimedia content 110A. The language model may include the audio content analyzer 420, the emotion classifier 422, and the search recommender 424.
[0116] The audio content analyzer 420 may extract meaningful insights from the audio content of the multimedia content 110A. The extraction may include at least one of transcription of speech to text, identification of speakers, detection of emotions, or analysis of the context of conversations. For instance, the audio content analyzer 420 may enable applications like voice recognition, sentiment analysis, and real-time audiobased information retrieval by integration of linguistic analysis with auditory input. Forexample, the audio content analyzer 420 may process spoken dialogue and background sounds.
[0117] The emotion classifier 422 may identify and categorize emotions expressed (such as joy, sadness, anger, or surprise) in the plurality of segments of the multimedia content 110A, such as text or speech. For instance, the emotion classifier 422 may analyze linguistic patterns, tone, and context, to enable applications like sentiment analysis, emotional profiling, and tailored responses to enhance human-computer interaction and understand user intent. For example, the emotion classifier 422 may assess the emotional tone of the multimedia content 110A or the moderated media content.
[0118] The search recommender 424 of the second electronic device 104 may generate content recommendation based on at least one of the profile information, the user input, or the context information. Based on analysis of patterns, intent, and relevant data, the search recommender 424 may enhance search efficiency, thus may optimize the content recommendation to guide the user 114A toward the most relevant moderated media content. For example, the search recommender 424 may work in conjunction with the content recommendation engine 406 to enhance content discovery. The search recommender 424 may analyze user search patterns and viewing history to refine content suggestions.
[0119] The first electronic device 102 may also allow for user input to modify the content moderation model 102A. For instance, the user input may include a feedback on the moderated media content or the effectiveness of the applied content moderation model 102A. The first electronic device 102 may use the user input to update and refine thecontent moderation model 102A, to improve the accuracy and alignment with user preferences over time.
[0120] By integration of these various modules and functionalities, the content moderation model 102A may offer a comprehensive, adaptive, and user-centric approach to content filtering and recommendation. The first electronic device 102 may continuously analyze the multimedia content 110A, the user behavior, and environmental factors to determine a tailored viewing experience that balances entertainment value with appropriate content moderation.
[0121] FIG. 5 is a flowchart illustrating an example method for multimedia content moderation based on content context, region-specific profanity, and user profile, in accordance with an embodiment of the disclosure. FIG. 5 is described in conjunction with elements from FIG. 1 , FIG. 2A, FIG. 2B, FIG. 3, and FIG. 4. With reference to FIG. 5, there is shown an exemplary flowchart 500. The flowchart 500 includes steps 502 to 520 for moderation of the multimedia content 110A based on content context, region-specific profanity, and user profile. The various operations of the flowchart 500 may be performed by any computing device, such as, the first electronic device 102 of FIG. 1 or the circuitry 202A of FIG. 2A. The flowchart starts from 502 and proceeds to 504.
[0122] At 504, multimedia content and content metadata associated with a plurality of segments of the multimedia content may be received from a second electronic device. The circuitry 202A may be configured to receive the multimedia content 110A and content metadata 110B associated with the multimedia content 110A from the second electronic device 104 via the communication network 112. This network interface 208A of the first electronic device 102 may establish a connection with the second electronic device 104to facilitate a data transfer including reception of the multimedia content 110A and the content metadata 11 OB. The multimedia content 110A may include at least one of broadcast content, video-on-demand content, or streaming content. The content metadata 110B may include information such as scene descriptions, dialogue transcripts, and audio characteristics for each segment of the multimedia content 110A. Details related to the reception of multimedia content and content metadata are described further, for example, in FIG. 3.
[0123] At 506, context information associated with each segment of the plurality of segments may be determined, based on the content metadata. The circuitry 202A may be configured to determine the context information associated with each segment of the plurality of segments, based on the content metadata 11 OB. The circuitry 202A may analyze the content metadata 110B using the content moderation model 102A to extract relevant context information for each segment. This process may involve leveraging the capabilities of the Vision Foundation Model (VFM) 104A, Audio Language Model (ALM) 104C, and Neural Language Model (NLM) 104B to perform comprehensive analysis of visual, audio, and textual elements within the content. For instance, the context information may include factors such as scene type, dialogue content, emotional tone, and presence of potentially sensitive topics. The determination of context information is elaborated further, for example, in FIG. 3 and FIG. 4.
[0124] At 508, region-specific profanity information associated with the multimedia content may be received from the second electronic device. The circuitry 202A may be configured to receive the region-specific profanity information associated with the multimedia content 110A from the second electronic device 104. The second electronicdevice 104 retrieve the region-specific profanity information from the database 108 through the server 106 and transmit the retrieved region-specific profanity information to the first electronic device 102. The use of region-specific profanity information for content moderation may ensure that local cultural norms and sensitivities may be taken into account for the content moderation. For example, certain words or phrases may be considered more offensive in some regions than others, and this information allows for more nuanced content filtering. The reception of region-specific profanity information is described in more detail, for example, in FIG. 3 and FIG. 4.
[0125] At 510, sensor data associated with a user present in a predetermined proximity of the first electronic device may be received from a set of sensors. The circuitry 202A may be configured to receive, from the set of sensors 102B, the sensor data associated with the user 114A present in a predetermined proximity (e.g., 5 meters) of the first electronic device 102. The circuitry 202A may be configured to collect data from the set of sensors 102B, which may include cameras, microphones, and biometric sensors. The sensor data may extract user profile information associated with at least one of: the viewer's presence, age, emotional state, and other relevant characteristics. For instance, a camera module within the set of sensors 102B may detect the presence and estimate the age of viewers, while an audio sensor might analyze ambient sounds to determine if multiple users are present. The collection of sensor data is further elaborated, for example, in FIG. 3 and FIG. 4.
[0126] At 512, profile information associated with the user may be determined based on the sensor data. The circuitry 202A may be configured to determine the profile information associated with the user 114A based on the sensor data. The circuitry 202Amay be configured to process the collected sensor data to construct a detailed user profile. This profile may include attributes such as the viewer's age, viewing preferences, emotional responses to different content types, and potentially even cultural background inferred from language or accent detection. For example, facial recognition algorithms may be used to estimate the viewer's age, while emotion recognition techniques may gauge their reactions to content. The determination of user profile information is described in more detail, for example, in FIG. 3 and FIG. 4.
[0127] At 514, a content moderation model may be applied on the plurality of segments, based on the context information, the region-specific profanity information, and the profile information. The circuitry 202A may be configured to apply the content moderation model 102A on the plurality of segments based on the context information, the region-specific profanity information, and the profile information. The circuitry 202A may analyze and filter the content segments. Multiple factors may be simultaneously factored to generate a nuanced and personalized content moderation experience. For instance, if the user profile indicates a teenager, the circuitry 202A may apply content moderation such that mild profanity may be allowed but segments with extreme violence may be flagged. The application of the content moderation model 102A is further elaborated, for example, in FIG. 3 and FIG. 4.
[0128] At 516, a set of unsuitable segments for the user may be selected from the plurality of segments, based on the application of the content moderation model. The circuitry 202A may be configured to select, from the plurality of segments, the set of unsuitable segments for the user 114A, based on the application of the content moderation model 102A. The circuitry 202A may be configured to identify segments thatmay deemed inappropriate for the specific users (e.g., the user 114A) based on their profile and the context of the content. This selection process may involve flagging segments including excessive violence, inappropriate language, or thematic elements that are not suitable for the viewer's age or preferences. The selection of unsuitable segments is described in more detail, for example, in FIG. 3 and FIG. 4.
[0129] At 518, a moderated media content may be determined based on the multimedia content and the set of unsuitable segments. The circuitry 202A may determine the moderated media content based on the multimedia content 110A and the set of unsuitable segments. The circuitry 202A may be configured to create a modified version of the original content that addresses the identified unsuitable segments. This may involve various techniques such as pixelation of graphic visual elements, replacement inappropriate audio, or seamless substitution of alternative scenes or footage. For example, in a video containing violent scenes, the circuitry 202A may apply visual effects to obscure graphic content while narrative continuity may be maintained. The determination of moderated media content is further elaborated, for example, in FIG. 3 and FIG. 4.
[0130] At 520, the user device may be controlled. The circuitry 202A may control the user device 114 to render the moderated media content. The circuitry 202A may be configured to manage the playback of the moderated content on the user device 114. The control of user device 114 may ensure that the viewer receives an appropriately filtered version of the content that aligns with user profile information and regional standards. The moderated content rendering may also include dynamic adjustment of the moderation in real-time based on changes in the viewing environment or user state detected by the setof sensors 102B. The moderated content rendering is described in more detail, for example, in FIG. 3 and FIG. 4. Control may pass to end.
[0131] Although the exemplary method is illustrated as discrete operations, such as 504, 506, 508, 510, 512, 514, 516, 518, and 520, the disclosure is not so limited. Accordingly, in certain embodiments, such discrete operations may be further divided into additional operations, combined into fewer operations, or eliminated, depending on the implementation without detracting from the essence of the disclosed embodiments.
[0132] Various embodiments of the disclosure may provide a non-transitory computer- readable medium and / or storage medium having stored thereon, computer-executable instructions by a machine and / or a computer to operate an electronic device (for example, the first electronic device 102 of FIG. 1). Such instructions may cause the first electronic device 102 to perform operations that may include receipt of multimedia content (e.g., the multimedia content 110A) and content metadata (e.g., the content metadata 110B) associated with a plurality of segments of the multimedia content 110A from a second electronic device (e.g., the second electronic device 104). The operations may further include determination of context information associated with each segment of the plurality of segments, based on the content metadata 110B. The operations may further include receipt of region-specific profanity information associated with the multimedia content 110A from the second electronic device 104. The operations may further include receipt of sensor data, from a set of sensors (e.g., the set of sensors 102B), associated with a user (e.g., the user 114A) present in a predetermined proximity of the first electronic device 102. The operations may further include determination of profile information associated with the user 114A based on the sensor data. The operations may furtherinclude application of a content moderation model (e.g., the content moderation model 102A) on the plurality of segments, based on the context information, the region-specific profanity information, and the profile information. The operations may further include selection of a set of unsuitable segments for the user 114A from the plurality of segments, based on the application of the content moderation model 102A. The operations may further include determination of a moderated media content based on the multimedia content 110A and the set of unsuitable segments. The operations may further include control of the user device 114 to render the moderated media content.
[0133] Exemplary aspects of the disclosure may provide an electronic device (such as, the first electronic device 102 of FIG. 1 ) that includes circuitry (such as, the circuitry 202A). The circuitry 202A may be configured to receive multimedia content (e.g., the multimedia content 110A) and content metadata (e.g., the content metadata 110B) associated with a plurality of segments of the multimedia content 110A from a second electronic device (e.g., the second electronic device 104). The circuitry 202A may be further configured to determine context information associated with each segment of the plurality of segments, based on the content metadata 11 OB. The circuitry 202A may be further configured to receive region-specific profanity information associated with the multimedia content 110A from the second electronic device 104. The circuitry 202A may be further configured to receive sensor data, from the set of sensors 102B, associated with a user (e.g., the user 114A) present in a predetermined proximity of the first electronic device 102. The circuitry 202A may be further configured to determine profile information associated with the user 114A based on the sensor data. The circuitry 202A may be further configured to apply a content moderation model (e.g., the content moderation model 102A) on the plurality ofsegments, based on the context information, the region-specific profanity information, and the profile information. The circuitry 202A may be further configured to select a set of unsuitable segments for the user 114A from the plurality of segments, based on the application of the content moderation model 102A. The circuitry 202A may be further configured to determine a moderated media content based on the multimedia content 110A and the set of unsuitable segments. The circuitry 202A may be further configured to control the user device 114 to render the moderated media content.
[0134] The circuitry 202A may be further configured to determine an age of the user 114A based on the sensor data and determine the profile information based on the age of the user 114A. The set of sensors 102B may include at least one of: a camera, a microphone, or a biometric sensor. The circuitry 202A may be further configured to detect a change in the sensor data, update the profile information based on the change in the sensor data, and dynamically update the moderated media content based on the updated profile information. The circuitry 202A may be further configured to generate a content recommendation based on the profile information and control the user device 114 to render the content recommendation.
[0135] The circuitry 202A may be further configured to determine a geographic location of the first electronic device 102 and receive the region-specific profanity information based on the determined geographic location. The circuitry 202 may be further configured to mask or replace the set of unsuitable segments in the moderated media content.
[0136] The circuitry 202A may be configured to receive user input to modify the content moderation model. The circuitry 202A may be further configured to update the content moderation model 102A based on the user input. The circuitry 202A may be configuredto generate the content moderation report based on the set of unsuitable segments. The circuitry 202A may be further configured to transmit the content moderation report to the second electronic device 104. The multimedia content 110 may include at least one of broadcast content, video-on-demand content, or streaming content.
[0137] A circuitry (e.g., the circuitry 202B) of the second electronic device 104 may be configured to apply a vision foundation model (VFM) (e.g., the VFM 104A) to analyze visual content of the multimedia content 110A. The circuitry 202B may be further configured to determine the context information based on the analysis of the visual content and transmit the context information to the first electronic device 102. The circuitry 202B may be further configured to apply a language model (LM) to analyze audio content of the multimedia content. The circuitry 202B may be further configured to determine the context information based on the analysis of the audio content and transmit the context information to the first electronic device 102. In an embodiment, the LM may correspond to at least one of a neural language model (e.g., the NLM 104B) or an audio language model (e.g., the ALM 104C).
[0138] The circuitry 202B of the second electronic device 104 may be configured to segment the multimedia content 110A into the plurality of segments. The circuitry 202B may be further configured to generate the content metadata 110B for each segment of the plurality of segments. The circuitry 202B may be further configured to transmit the multimedia content 110A and the content metadata 110B to the first electronic device102. The circuitry 202B may be configured to retrieve the region-specific profanity information from a database (e.g., the database 108). The circuitry 202B may be furtherconfigured to transmit the region-specific profanity information to the first electronic device 102.
[0139] The present disclosure may also be positioned in a computer program product, which include all the features that enable the implementation of the methods described herein, and which when loaded in a computer system is able to conduct these methods. Computer program, in the present context, means any expression, in any language, code or notation, of a set of instructions intended to cause a system with information processing capability to perform a particular function either directly, or after either or both of the following: a) conversion to another language, code or notation; b) reproduction in a different material form.
[0140] While the present disclosure is described with reference to certain embodiments, it will be understood by those skilled in the art that various changes may be made, and equivalents may be substituted without departure from the scope of the present disclosure. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the present disclosure without departure from its scope. Therefore, it is intended that the present disclosure is not limited to the embodiment disclosed, but that the present disclosure will include all embodiments that fall within the scope of the appended claims.
Claims
CLAIMSWhat is claimed is:1 . An electronic device, comprising: circuitry configured to: receive multimedia content and content metadata associated with a plurality of segments of the multimedia content from a second electronic device; determine context information associated with each segment of the plurality of segments, based on the content metadata; receive region-specific profanity information associated with the multimedia content from the second electronic device; receive, from a set of sensors, sensor data associated with a user present in a predetermined proximity of the first electronic device; determine profile information associated with the user based on the sensor data; apply a content moderation model on the plurality of segments, based on the context information, the region-specific profanity information, and the profile information; select, from the plurality of segments, a set of unsuitable segments for the user, based on the application of the content moderation model; determine a moderated media content based on the multimedia content and the set of unsuitable segments; and control a user device to render the moderated media content.
2. The first electronic device according to claim 1 , wherein the circuitry is further configured to: determine an age of the user based on the sensor data; and determine the profile information based on the age of the user.
3. The first electronic device according to claim 1 , wherein the set of sensors comprises at least one of: a camera, a microphone, or a biometric sensor.
4. The first electronic device according to claim 1 , wherein the circuitry is further configured to: detect a change in the sensor data; update the profile information based on the change in the sensor data; and dynamically update the moderated media content based on the updated profile information.
5. The first electronic device according to claim 1 , wherein the circuitry is further configured to: generate a content recommendation based on the profile information; and control the user device 114 to render the content recommendation.
6. The first electronic device according to claim 1 , wherein the second electronic device is configured to:apply a vision foundation model to analyze visual content of the multimedia content; determine the context information based on the analysis of the visual content; and transmit the context information to the first electronic device.
7. The first electronic device according to claim 1 , wherein the second electronic device is configured to: apply a language model to analyze audio content of the multimedia content; determine the context information based on the analysis of the audio content; and transmit the context information to the first electronic device.
8. The first electronic device according to claim 7, wherein the language model corresponds to at least one of a neural language model or an audio language model.
9. The first electronic device according to claim 1 , wherein the circuitry is further configured to: determine a geographic location of the first electronic device; and receive the region-specific profanity information based on the determined geographic location.
10. The first electronic device according to claim 1 , wherein the circuitry is further configured to mask or replace the set of unsuitable segments in the moderated media content.11 . The first electronic device according to claim 1 , wherein the second electronic device is configured to: segment the multimedia content into the plurality of segments; generate the content metadata for each segment of the plurality of segments; and transmit the multimedia content and the content metadata to the first electronic device.
12. The first electronic device according to claim 1 , wherein the second electronic device is configured to: retrieve the region-specific profanity information from a database; and transmit the region-specific profanity information to the first electronic device.
13. The first electronic device according to claim 1 , wherein the circuitry is further configured to: receive user input to modify the content moderation model; and update the content moderation model based on the user input.
14. The first electronic device according to claim 1 , wherein the circuitry is further configured to: generate a content moderation report based on the set of unsuitable segments; and transmit the content moderation report to the second electronic device.
15. The first electronic device according to claim 1 , wherein the multimedia content comprises at least one of broadcast content, video-on-demand content, or streaming content.
16. A method, comprising: in a first electronic device: receiving multimedia content and content metadata associated with a plurality of segments of the multimedia content from a second electronic device; determining context information associated with each segment of the plurality of segments, based on the content metadata; receiving region-specific profanity information associated with the multimedia content from the second electronic device; receiving, from a set of sensors, sensor data associated with a user present in a predetermined proximity of the first electronic device; determining profile information associated with the user based on the sensor data;applying a content moderation model on the plurality of segments, based on the context information, the region-specific profanity information, and the profile information; selecting, from the plurality of segments, a set of unsuitable segments for the user, based on the application of the content moderation model; determining a moderated media content based on the multimedia content and the set of unsuitable segments; and controlling a user device to render the moderated media content.
17. The method according to claim 16, further comprising: determining an age of the user based on the sensor data; and determining the profile information based on the age of the user.
18. The method according to claim 16, wherein the set of sensors comprises at least one of: a camera, a microphone, or a biometric sensor.
19. The method according to claim 16, further comprising: detecting a change in the sensor data; updating the profile information based on the change in the sensor data; and dynamically updating the moderated media content based on the updated profile information.
0. A non-transitory computer-readable medium having stored thereon, computerexecutable instructions that when executed by a first electronic device, causes the first electronic device to execute operations, the operations comprising: receiving multimedia content and content metadata associated with a plurality of segments of the multimedia content from a second electronic device; determining context information associated with each segment of the plurality of segments, based on the content metadata; receiving region-specific profanity information associated with the multimedia content from the second electronic device; receiving, from a set of sensors, sensor data associated with a user present in a predetermined proximity of the first electronic device; determining profile information associated with the user based on the sensor data; applying a content moderation model on the plurality of segments, based on the context information, the region-specific profanity information, and the profile information; selecting, from the plurality of segments, a set of unsuitable segments for the user, based on the application of the content moderation model; determining a moderated media content based on the multimedia content and the set of unsuitable segments; and controlling a user device to render the moderated media content.
Citation Information
Patent Citations
Video player censor settings
US20150067717A1
Automatic Content Presentation Adaptation Based on Audience
US20200021888A1
Content filtering in media playing devices
US20210329338A1
Systems and methods for presenting content simultaneously in different forms based on parental control settings
US20240107100A1
Modifying Existing Content Based on Target Audience
US20240179374A1