Multi-modal content security analysis method, device, equipment, medium and product

By collecting multimodal content data on the Internet TV side and utilizing a risk detection model based on a gradient boosting framework, the problems of low real-time performance and complex rule configuration in existing technologies are solved, enabling real-time personalized analysis of Internet TV content security.

CN121151633APending Publication Date: 2025-12-16CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411550191.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing multimodal content analysis solutions suffer from low real-time performance and complex review rule configuration, failing to meet the real-time requirements of Internet TV content security, and making it difficult for users to customize rules that take effect in real time for personalized scenarios.

Method used

Multimodal content data is collected on edge devices, and personalized rules are generated using multimodal content recognition models and scene discriminators. These rules are then combined with a risk detection model based on a gradient boosting framework for security analysis, reducing reliance on network bandwidth and computing resources and enabling real-time risk detection.

Benefits of technology

It improves the real-time and portability of content security analysis, reduces network congestion and transmission latency, adapts to personalized scenario needs, and enables instant risk detection and rule configuration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121151633A_ABST
    Figure CN121151633A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal content security analysis method and device, equipment, a medium and a product, and belongs to the technical field of artificial intelligence, and the method comprises the steps: responding to a data collection task issued by an end-side user, and collecting multi-modal content data; determining a multi-modal feature corresponding to the multi-modal content data according to a multi-modal content identification model; determining a target scene corresponding to the multi-modal feature based on a scene discriminator, wherein the target scene is used for generating a personalized rule with a preset rule base and / or a user-defined rule; and inputting the multi-modal features and the personalized rules into a pre-trained risk detection model to obtain a safety analysis result output by the risk detection model, the risk detection model being a model based on a gradient lifting framework. According to the method, the multi-modal data acquisition and analysis process is transferred from the cloud to the edge, the real-time performance is improved, the target scene and the preset rule base and / or the user-defined rule generate the personalized rule, and the rule which takes effect in real time can be customized for the personalized scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a multimodal content security analysis method, apparatus, equipment, medium and product. Background Technology

[0002] With the rapid development of information technology and the widespread application of technologies such as 5G and artificial intelligence, the transmission speed and viewing experience of Internet TVs have been further improved, and consumers' demands for televisions are constantly upgrading, such as high-definition picture quality, massive media resources, and intelligent recommendations. As a home product, Internet TVs target a wide range of people, covering all age groups, including the elderly and teenagers. Therefore, ensuring the security of Internet TV content is crucial. As a high-performance, high-definition, and intelligent Internet TV product, it contains a massive and real-time updated Electronic Program Guide (EPG) pages, media resources, live and on-demand resources, involving multimodal content data such as text, images, video, and audio.

[0003] Existing multimodal content analysis solutions typically require data transmission to the cloud for processing. When there are too many content resources under test, this can cause network congestion and transmission delays, impacting the real-time performance and immediacy of monitoring. However, in real-world monitoring scenarios, the review of online content security demands high real-time performance; insufficient real-time performance may lead to the continued escalation of content security risks. Review rules are often fixed and configured on the server side, making configuration complex and requiring dedicated personnel for maintenance. Users find it difficult to customize rules that take effect in real-time for individual scenarios, resulting in delays in review requirements.

[0004] In summary, existing multimodal content analysis solutions suffer from problems such as low real-time performance and complex configuration of review rules. Summary of the Invention

[0005] This application provides a multimodal content security analysis method, apparatus, equipment, medium, and product, which solves the shortcomings of existing technologies such as low real-time performance and complex configuration of review rules.

[0006] This application provides a multimodal content security analysis method, including: In response to data collection tasks issued by users on the client side, collect multimodal content data; The multimodal content data is input into a preset multimodal content recognition model, which is used to determine the multimodal features corresponding to the multimodal content data. The target scene corresponding to the multimodal features is determined based on a preset scene discriminator, and the target scene is used to generate personalized rules with a preset rule base and / or custom rules; The multimodal features and the personalized rules are input into a pre-trained risk detection model to obtain the security analysis results output by the risk detection model, which is a model based on the gradient boosting framework.

[0007] As an example, the preset rule base and / or the custom rules include risk rules corresponding to each scenario and multi-level risk labels corresponding to each risk rule. Correspondingly, the security analysis results include risk rule analysis results, risk label analysis results, and risk confidence levels.

[0008] As one embodiment, the step of collecting multimodal content data in response to a data collection task issued by the user on the client side includes: In response to a data acquisition task issued by a user on the terminal side through a terminal control device, multimodal content data is acquired based on the on-demand media asset acquisition method, the on-demand video acquisition method, and the Electronic Program Guide (EPG) acquisition method. The multimodal content data includes video, audio, images, and text.

[0009] This application also provides a multimodal content security analysis device for implementing the multimodal content security analysis method. The device includes a receiving control circuit, a data acquisition card circuit, and an acceleration processor. The receiving control circuit and the data acquisition card circuit are respectively connected to the acceleration processor. The receiving control circuit is used to receive data acquisition tasks sent by users on the receiving end side. The acquisition card circuit is used to acquire multimodal content data according to the data acquisition task; The acceleration processor is used to input the multimodal content data into a preset multimodal content recognition model, which is used to determine the multimodal features corresponding to the multimodal content data; determine the target scene corresponding to the multimodal features based on a preset scene discriminator, which is used to generate personalized rules with a preset rule base and / or custom rules; input the multimodal features and the personalized rules into a pre-trained risk detection model to obtain the security analysis results output by the risk detection model, which is a model based on a gradient boosting framework.

[0010] As one embodiment, the device further includes a display control circuit connected to the acceleration processor. The display control circuit is connected to a touch screen and / or an external display screen. The display control circuit is also used to receive data acquisition tasks issued by the end-side user based on the touch screen; and to send the security analysis results to the external display screen so that the external display screen can display the security analysis results.

[0011] As an example, the acceleration processor is also used to update or synchronize the risk detection model and the preset rule base and / or custom rules.

[0012] As an example, the acceleration processor is also used to send the security analysis results to a preset cloud platform.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multimodal content security analysis method as described above.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multimodal content security analysis method as described above.

[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the multimodal content security analysis method as described above.

[0016] The multimodal content security analysis method, apparatus, device, medium, and product provided in this application collect multimodal content data in response to data collection tasks issued by end-side users via edge devices; input the multimodal content data into a preset multimodal content recognition model, which is used to determine the multimodal features corresponding to the multimodal content data; determine the target scene corresponding to the multimodal features based on a preset scene discriminator, which is used to generate personalized rules with a preset rule base and / or custom rules; input the multimodal features and the personalized rules into a pre-trained risk detection model to obtain the security analysis results output by the risk detection model, wherein the risk detection model is a model based on a gradient boosting framework. Moving the multimodal data collection and analysis process from the cloud to the edge effectively reduces reliance on network bandwidth and computing resources, solves network limitations, avoids network congestion and transmission delays, and improves real-time performance and portability. Personalized rules are generated from the target scenario and the preset rule base and / or custom rules, and then risk detection is performed based on the risk detection model constructed by the gradient boosting framework. This solves the problem that users have difficulty in customizing rules that take effect in real time for personalized scenarios, which causes delays in review requirements. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is one of the flowcharts of the multimodal content security analysis method provided in this application.

[0019] Figure 2 This is the second flowchart of the multimodal content security analysis method provided in this application.

[0020] Figure 3 This is one of the structural schematic diagrams of the multimodal content security analysis device provided in this application.

[0021] Figure 4 This is the second schematic diagram of the multimodal content security analysis device provided in this application.

[0022] Figure 5 This is a schematic diagram of the workflow of the multimodal content security analysis device provided in this application.

[0023] Figure 6 This is a schematic diagram of the structure of the multimodal content security analysis system provided in this application.

[0024] Figure 7 This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] It should be noted that all actions involving the acquisition of signals, information, or data in this application are carried out in compliance with the relevant data protection laws and regulations of the locality and with authorization from the owner of the relevant device.

[0027] The existing review of Internet TV on-demand and EPG resources is usually conducted manually or by machine by the licensee before the content is injected. There is a problem of inconsistent review standards, and the content security review standards are dynamic. Therefore, it is necessary to carry out routine security monitoring of the content that has been launched.

[0028] Existing technologies suffer from problems such as inconsistent review standards, low real-time performance, inability to fully utilize multimodal data relationships, and complex review rule configuration. To address at least one of these issues, this application provides a multimodal content security analysis method, apparatus, equipment, medium, and product that can be applied to business scenarios such as Internet TV content security assurance. The following is a detailed description in conjunction with the accompanying drawings.

[0029] Figure 1 This is one of the flowcharts illustrating the multimodal content security analysis method provided in this application. Figure 2 This is the second flowchart of the multimodal content security analysis method provided in this application, as shown below. Figure 1 and Figure 2 As shown, this application provides a multimodal content security analysis method, including steps S100-S400.

[0030] Step S100: In response to the data collection task issued by the user on the terminal side, collect multimodal content data.

[0031] Currently, broadcast control operators provide set-top boxes for two types of services: Over-The-Top (OTT) video services and Internet Protocol Television (IPTV). IPTV set-top boxes typically transmit content data via a dedicated network channel (private network), while OTT set-top boxes, although using the public network, also face cross-regional network connectivity issues. Therefore, monitoring set-top box content usually involves remotely issuing tasks to collect data locally, then synchronizing media asset information and audio / video streams to a monitoring server for monitoring and processing using a high-performance GPU server. This makes it difficult for end-users to intuitively perceive the task status, undoubtedly increasing operational complexity. When multiple acquisition devices operate simultaneously, it often creates bandwidth pressure, increasing network transmission costs and impacting review efficiency.

[0032] To address this issue, this application issues data collection tasks on the client side and executes them on the edge side, collecting multimodal content data from media assets, on-demand content, and EPG pages. Specifically, it extracts video, audio, images, and text from the collected on-demand media assets, on-demand videos, and EPG images through methods such as audio-video separation, frame extraction, frame cropping, and data extraction.

[0033] Step S200: Input the multimodal content data into a preset multimodal content recognition model, wherein the multimodal content recognition model is used to determine the multimodal features corresponding to the multimodal content data.

[0034] Optionally, multimodal content recognition models include flag / logo detection models, face detection models, risk recognition models, behavioral risk recognition models, object recognition models, keyword detection models, optical character recognition (OCR) models, automatic speech recognition (ASR) models, semantic analysis models, sentiment analysis models, voiceprint recognition models, and tone recognition models. Voice data is processed into text data using ASR while simultaneously undergoing voiceprint and tone recognition; video data is obtained by frame capture; image data is processed through OCR for text recognition while also performing lightweight model recognition for risk, scene, behavior, face, and logo detection; and text data undergoes keyword detection, semantic analysis, and sentiment analysis.

[0035] Step S300: Determine the target scene corresponding to the multimodal features based on a preset scene discriminator. The target scene is used to generate personalized rules with a preset rule base and / or custom rules.

[0036] A scene discriminator is a neural network model designed to classify input images, determining which scene or category the image belongs to. It can be combined with other input modalities, such as combining text descriptions with images, to determine whether the image matches a given text description.

[0037] This application pre-classifies multiple scenarios based on the content categories of media assets, on-demand content, and EPG pages, such as news, movies, children's programs, sports competitions, reality shows, and documentaries. In other embodiments, scenario categories can also be determined according to other classification methods, and this application does not limit this.

[0038] As an example, the preset rule base and / or the custom rules include risk rules corresponding to each scenario and multi-level risk labels corresponding to each risk rule. Correspondingly, the security analysis results include risk rule analysis results, risk label analysis results, and risk confidence levels.

[0039] After identifying the scenarios, risk rules need to be configured for each scenario. These risk rules include general risk rules, such as prohibiting violations of laws and regulations. In addition to these general rules, specific risk rules are also set, such as prohibiting content that has a negative impact on children in children's program scenarios. The risk rules for all scenarios are then combined to form a preset rule library. Custom rules refer to risk rules that users define for specific scenarios.

[0040] The multi-level risk label includes multiple primary labels and multiple secondary labels. Secondary labels are sub-labels of primary labels. Tertiary labels can also be set under secondary labels, and so on. Primary labels can be set according to actual needs, and this application does not limit them.

[0041] Step S400: Input the multimodal features and the personalized rules into the pre-trained risk detection model to obtain the security analysis results output by the risk detection model. The risk detection model is a model based on the gradient boosting framework.

[0042] Specifically, the risk detection model is the Scene-based eXtreme Gradient Boosting Machine for Content Security analysis (SXGBMCS). The extreme gradient boosting machine is an efficient machine learning algorithm that is widely used in classification and regression problems. Based on the principle of gradient boosting, it creates a strong learner by combining multiple weak learners (usually decision trees).

[0043] The SXGBMCS model in this application is based on a gradient boosting framework. It integrates multiple weak learners, uses a fixed scene as the first feature of the algorithm, and optimizes hyperparameters through a genetic algorithm to avoid the negative impact of manual parameter tuning. It uses scene-risk as the primary evaluation metric to construct a highly reliable prediction model. The SXGBMCS model is adaptable to multimodal data and has good model interpretability. This application uses the output of the multimodal content recognition model and personalized rules as input to output the final security analysis result. In other embodiments, the model decision-making process can be visualized.

[0044] Understandably, this application moves the multimodal data collection and analysis process from the cloud to the edge, effectively reducing reliance on network bandwidth and computing resources, solving network limitation problems, avoiding network congestion and transmission delays, and improving real-time performance and portability. Personalized rules are generated from the target scenario and the preset rule base and / or custom rules, and then risk detection is performed based on the risk detection model constructed by the gradient boosting framework. This solves the problem that users have difficulty in customizing rules that take effect in real time for personalized scenarios, which causes a delay in review requirements.

[0045] Based on the above embodiments, as an optional embodiment, the step of collecting multimodal content data in response to a data collection task issued by the end-side user includes: in response to a data collection task issued by the end-side user through the end-side control device, collecting multimodal content data based on the on-demand media asset collection method, the on-demand video collection method, and the Electronic Program Guide (EPG) collection method, wherein the multimodal content data includes video, audio, images, and text.

[0046] Optionally, end-users can issue data collection tasks via touchscreen or Bluetooth remote control, and the collection methods include media asset collection, video-on-demand collection, and EPG collection.

[0047] Media asset acquisition: Obtain full media asset data through packet capture analysis and parsing interfaces, including but not limited to text and image data such as media asset type, name, poster, region, introduction, director, actors, and playback links.

[0048] Video on demand acquisition: Parse the video on demand links obtained from media asset acquisition, and then obtain audio and video stream data.

[0049] EPG Acquisition: By connecting to a set-top box, the video and audio are directly transmitted to the acquisition device to obtain EPG image data in real time.

[0050] Understandably, this application migrates the Internet TV content analysis process from the cloud to the edge via edge computing. The entire process can be operated independently at the edge, with data collection tasks issued by the edge user. This avoids limitations imposed by specific network scenarios, reduces reliance on network bandwidth, solves network limitation problems, avoids network congestion and transmission delays, improves real-time performance and immediacy, and at the same time reduces the demand for computing resources and improves performance efficiency.

[0051] The multimodal content security analysis apparatus provided in this application is described below. The multimodal content security analysis apparatus described below and the multimodal content security analysis method described above can be referred to in correspondence.

[0052] Figure 3 This is a timing diagram of the multimodal content security analysis device provided in this application, such as... Figure 3 As shown, this application also provides a multimodal content security analysis device for implementing the multimodal content security analysis method, including a receiving control circuit 310, a data acquisition card circuit 320, and an acceleration processor 330, wherein the receiving control circuit 310 and the data acquisition card circuit 320 are respectively connected to the acceleration processor 330.

[0053] The receiving control circuit 310 is used to receive data acquisition tasks issued by the user on the receiving end; specifically, the receiving control circuit 310 can be an infrared / Bluetooth receiver, and the user on the receiving end is equipped with an infrared / Bluetooth transmitter.

[0054] The acquisition card circuit 320 is used to acquire multimodal content data according to the data acquisition task. The acquisition card circuit 320 is responsible for acquiring Internet TV EPG pages, media assets, and on-demand resources, and converting them into multimodal data such as text, images, and audio, which are then transmitted to the data analysis module. Optionally, the acquisition card circuit 320 can be an HDMI acquisition card.

[0055] The acceleration processor 330 is used to input the multimodal content data into a preset multimodal content recognition model, which is used to determine the multimodal features corresponding to the multimodal content data; determine the target scene corresponding to the multimodal features based on a preset scene discriminator, which is used to generate personalized rules with a preset rule base and / or custom rules; input the multimodal features and the personalized rules into a pre-trained risk detection model to obtain the security analysis results output by the risk detection model, which is a model based on a gradient boosting framework.

[0056] The accelerator processor 330 integrates a scene discriminator and a content review model. It automatically generates alarms based on different scenes such as animated movies, sports competitions, and news channels, as well as model prediction results, and reports the alarms to the backend.

[0057] Media assets, on-demand, and EPG data are collected by the acquisition card circuit 320 and output to the accelerator processor 330. The accelerator processor 330 performs multimodal content security intelligent analysis, including face detection analysis of target persons, matching analysis of prohibited logos, and text risk analysis after OCR text recognition and ASR speech-to-text conversion. It automatically identifies Internet TV monitoring scenarios, including but not limited to animated movies, children's programs, sports competitions, news reports, advertisements, short videos, etc. Finally, it combines custom rules and uses the SXGBMCS model to output risk content analysis results such as first-level risk labels.

[0058] Existing content security analysis models lack adaptation for Internet TV monitoring scenarios and use high-performance servers that cannot meet the performance requirements of edge devices. Therefore, this application prunes and quantizes the risk detection model to improve its lightweightness, reduce its storage requirements and computational complexity, and utilizes the highly parallel architecture of the Neural Processor Unit (NPU) for hardware acceleration.

[0059] In one embodiment, this application utilizes text and poster data from media assets to construct a monitoring scene discriminator, which outputs candidate scenes according to priority to more accurately match the rule thresholds and models of different scenes. To meet personalized needs, this application supports end-users in adding custom keywords, logos, and combinations thereof as rules. Security analysis results can be output in real time, and anomaly information can be transmitted to the server to ensure timely response and problem backtracking.

[0060] Figure 4 This is the second structural schematic diagram of the multimodal content security analysis device provided in this application, as shown below. Figure 4 As shown in the figure, as an embodiment, the multimodal content security analysis device provided in this application further includes a display control circuit connected to the acceleration processor. The display control circuit is connected to a touch screen and / or an external display screen. The display control circuit is also used to receive data collection tasks issued by the end-side user based on the touch screen; and to send the security analysis results to the external display screen so that the external display screen can display the security analysis results.

[0061] The accelerator processor is also used to send the security analysis results to a preset cloud platform so that the cloud platform can centrally manage the security analysis results. The multimodal content security analysis device provided in this application is divided into two types: with a screen and without a screen. The device with a screen is operated via a touch screen, while the device without a screen can be operated via a Bluetooth remote control. It can be connected to smart terminals such as Internet TVs via an HDMI cable to display risk alarms detected on the edge in real time. Risk alarms are automatically generated when the risk confidence level is higher than the risk threshold.

[0062] Optionally, the display control circuit is also used to send current media asset information, real-time images, and multi-level tags, location, and confidence information such as real-time scenes, faces, logo libraries, and audio text to smart terminals such as internet TVs for display. This multi-angle, multi-dimensional information display helps users gain a deeper understanding of the risk detection results, improving the comprehensiveness and accuracy of risk detection. In another embodiment, it can also control the display of historical anomaly information, combined with progress adjustment and timestamp jump control, to quickly revisit problematic scenarios.

[0063] Understandably, the control and display module can be used to flexibly respond to different user needs, realize the real-time display of risk alarms detected on the end side, and overcome the shortcomings of existing security analysis devices in terms of visualization, alarm and backtracking.

[0064] As an example, the accelerator processor 330 is also used to update or synchronize the risk detection model and the preset rule base and / or custom rules, and is responsible for updating and synchronizing the rule base, feature base and model, configuring task rules and issuing tasks.

[0065] Specifically, the acceleration processor 330 is also used to control the set-top box in the EPG scenario, select monitoring tasks in media asset and on-demand scenarios, configure custom review rules, select real-time screen elements, adjust on-demand progress, and trigger SMS and email alarms.

[0066] Understandably, this application enables users to comprehensively and flexibly manage monitoring tasks and make personalized configurations as needed by setting up a task configuration module.

[0067] Figure 5 This is a schematic diagram of the workflow of the multimodal content security analysis device provided in this application, such as... Figure 5 As shown, based on the above embodiments, as an optional embodiment, the multimodal content security analysis device provided in this application can be integrated into a plug-and-play security analysis device. It is electrically connected to an internet TV via an internet TV interface, supporting control and alarm display of the internet TV. It also supports user-defined rule configuration, improving system flexibility compared to traditional static rules, adapting to dynamic and diverse content review needs, and responding promptly to emerging forms of illegal content. Compared to similar data collection devices that distribute tasks and collect data on the server side, this application can operate independently on the device side, featuring portability, visualization, and high operability. It can adapt to different network environments, thereby improving the user experience.

[0068] The multimodal content security analysis system provided in this application is described below. The multimodal content security analysis system described below can be referred to in correspondence with the multimodal content security analysis device described above.

[0069] Figure 6 This is a schematic diagram of the structure of the multimodal content security analysis system provided in this application, such as... Figure 6 As shown, this application also provides a multimodal content security analysis system, including the aforementioned multimodal content security analysis device, which has the technical features and effects corresponding to the multimodal content security analysis device, and will not be described in detail here.

[0070] As one embodiment, the multimodal content security analysis system further includes a cloud analysis device. The cloud analysis device is used to send data acquisition tasks to the multimodal content security analysis device, acquire multimodal content data fed back by the multimodal content security analysis device, input the multimodal content data into a preset multimodal content recognition model, the multimodal content recognition model being used to determine the multimodal features corresponding to the multimodal content data; determine the target scene corresponding to the multimodal features based on a preset scene discriminator, the target scene being used to generate personalized rules with a preset rule base and / or custom rules; input the multimodal features and the personalized rules into a pre-trained risk detection model to obtain the security analysis results output by the risk detection model, the risk detection model being a gradient boosting framework-based model.

[0071] When network conditions are good, the multimodal content security analysis device and the cloud analysis device can be controlled to perform combined security analysis.

[0072] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute a multimodal content security analysis method, which includes: collecting multimodal content data in response to a data collection task issued by a user on the edge; inputting the multimodal content data into a preset multimodal content recognition model, wherein the multimodal content recognition model is used to determine the multimodal features corresponding to the multimodal content data; determining the target scene corresponding to the multimodal features based on a preset scene discriminator, wherein the target scene is used to generate personalized rules with a preset rule base and / or custom rules; and inputting the multimodal features and the personalized rules into a pre-trained risk detection model to obtain the security analysis result output by the risk detection model, wherein the risk detection model is a model based on a gradient boosting framework.

[0073] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0074] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multimodal content security analysis method provided by the above methods. The method includes: collecting multimodal content data in response to a data collection task issued by a user on the terminal side; inputting the multimodal content data into a preset multimodal content recognition model, wherein the multimodal content recognition model is used to determine the multimodal features corresponding to the multimodal content data; determining the target scene corresponding to the multimodal features based on a preset scene discriminator, wherein the target scene is used to generate personalized rules with a preset rule base and / or custom rules; and inputting the multimodal features and the personalized rules into a pre-trained risk detection model to obtain the security analysis result output by the risk detection model, wherein the risk detection model is a model based on a gradient boosting framework.

[0075] In another aspect, this application also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program is implemented to perform the multimodal content security analysis method provided by the above methods. The method includes: collecting multimodal content data in response to a data collection task issued by a user on the edge; inputting the multimodal content data into a preset multimodal content recognition model, wherein the multimodal content recognition model is used to determine the multimodal features corresponding to the multimodal content data; determining a target scene corresponding to the multimodal features based on a preset scene discriminator, wherein the target scene is used to generate personalized rules with a preset rule base and / or custom rules; and inputting the multimodal features and the personalized rules into a pre-trained risk detection model to obtain a security analysis result output by the risk detection model, wherein the risk detection model is a model based on a gradient boosting framework.

[0076] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0077] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A multimodal content security analysis method, characterized in that, include: In response to data collection tasks issued by users on the client side, collect multimodal content data; The multimodal content data is input into a preset multimodal content recognition model, which is used to determine the multimodal features corresponding to the multimodal content data. The target scene corresponding to the multimodal features is determined based on a preset scene discriminator, and the target scene is used to generate personalized rules with a preset rule base and / or custom rules; The multimodal features and the personalized rules are input into a pre-trained risk detection model to obtain the security analysis results output by the risk detection model, which is a model based on the gradient boosting framework.

2. The multimodal content security analysis method according to claim 1, characterized in that, The preset rule base and / or the custom rules include risk rules corresponding to each scenario and multi-level risk labels corresponding to each risk rule. Correspondingly, the security analysis results include risk rule analysis results, risk label analysis results, and risk confidence levels.

3. The multimodal content security analysis method according to claim 1, characterized in that, The data collection task issued in response to the user on the client side collects multimodal content data, including: In response to a data acquisition task issued by a user on the terminal side through a terminal control device, multimodal content data is acquired based on the on-demand media asset acquisition method, the on-demand video acquisition method, and the Electronic Program Guide (EPG) acquisition method. The multimodal content data includes video, audio, images, and text.

4. A multimodal content security analysis device, characterized in that, The device is used to implement the multimodal content security analysis method according to any one of claims 1-3, the device includes a receiving control circuit, a data acquisition card circuit and an acceleration processor, wherein the receiving control circuit and the data acquisition card circuit are respectively connected to the acceleration processor; The receiving control circuit is used to receive data acquisition tasks sent by users on the receiving end side. The acquisition card circuit is used to acquire multimodal content data according to the data acquisition task; The acceleration processor is used to input the multimodal content data into a preset multimodal content recognition model, and the multimodal content recognition model is used to determine the multimodal features corresponding to the multimodal content data; The target scene corresponding to the multimodal features is determined based on a preset scene discriminator. The target scene is used to generate personalized rules with a preset rule base and / or custom rules. The multimodal features and the personalized rules are input into a pre-trained risk detection model to obtain the security analysis results output by the risk detection model. The risk detection model is a model based on a gradient boosting framework.

5. The multimodal content security analysis device according to claim 4, characterized in that, The device further includes a display control circuit connected to the acceleration processor. The display control circuit is connected to a touch screen and / or an external display screen. The display control circuit is also used to receive data acquisition tasks issued by the end-side user based on the touch screen; and to send the security analysis results to the external display screen so that the external display screen can display the security analysis results.

6. The multimodal content security analysis device according to claim 4 or 5, characterized in that, The accelerator is also used to update or synchronize the risk detection model and the preset rule base and / or custom rules.

7. The multimodal content security analysis device according to claim 6, characterized in that, The accelerator processor is also used to send the security analysis results to a preset cloud platform.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the multimodal content security analysis method as described in any one of claims 1 to 3.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the multimodal content security analysis method as described in any one of claims 1 to 3.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the multimodal content security analysis method as described in any one of claims 1 to 3.