Playing terminal and playing control method
By introducing a security management module to the playback terminal to safely identify and process the playback data, the problems of low efficiency and poor timeliness of playback content in the prior art are solved, and efficient security monitoring and processing in various scenarios are achieved.
Patent Information
- Application Number
- CN202411446472.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-16
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, the security monitoring efficiency of playback content is low, the timeliness is poor, and the application scenarios are limited, making it difficult to effectively prevent the playback of illegal content.
A playback terminal is designed, including a control module, a security management module, an image display module and a voice playback module. The security management module performs security identification of the data to be played. If the recognition result fails, the security processing will be performed before playing.
Real-time security identification and processing of the files to be played is realized, ensuring the security of the content to be played, improving the security monitoring efficiency, and suitable for a variety of playback scenarios, avoiding security accidents caused by illegal content playback.
Smart Images

Figure CN120223933A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of display technologies. More specifically, it relates to a playback terminal and a playback control method. Background Art
[0002] In the current information society, public display devices such as LCD TVs, advertising screens, information dissemination systems, etc. have become important media for information dissemination and are widely used in various public places such as enterprises and institutions, schools, hospitals, railway stations, shopping malls, etc. However, with the popularization of these display devices, the security issues of the playback content have become increasingly prominent. According to incomplete statistics, there are more than 3,000 accidents registered annually due to the playback of illegal content, seriously polluting the social atmosphere and likely causing adverse social impacts. For enterprises, once the display devices under their management play illegal content, it will constitute a major security accident, interfering with the normal operation of the enterprise and damaging the brand image of the enterprise.
[0003] Currently, for the security monitoring of the playback content of display devices, there are mainly the following two methods: one is manual monitoring, and the other is video annotation. However, both of these methods have obvious limitations. Manual monitoring is inefficient and lacks timeliness, while the method of video annotation is usually only applicable to specific playback scenarios, such as advertising playback scenarios, and is not applicable to the playback content of other scenarios. Summary of the Invention
[0004] The purpose of the present disclosure is to provide a playback terminal and a playback control method to solve the technical problems of low efficiency, poor timeliness, and limited application scenarios in the security monitoring of playback content in related technologies.
[0005] To achieve the above object, the present disclosure adopts the following technical solutions:
[0006] The first aspect of the present disclosure provides a playback terminal, including a control module, a security management module, an image display module, and a voice playback module, wherein the control module is connected to the security management module, the image display module, and the voice playback module;
[0007] The control module is configured to decode a file to be played to obtain data to be played, transmit the data to be played to the security management module and receive the security identification result feedback by the security management module, and when the security identification result is the first result, transmit the data to be played to the image display module and / or the voice playback module for playback, and when the security identification result is the second result, perform security processing on the data to be played and transmit the data to be played after security processing to the image display module and / or the voice playback module for playback, and the data to be played includes an image frame to be displayed and / or audio data;
[0008] The security management module is used to perform security identification on the data to be played and feedback the security identification result to the control module.
[0009] Optionally, the control module includes:
[0010] A first processing unit, configured to decode the file to be played and generate image data and / or audio data, segment the image data into image frames to be displayed and store them in a first storage unit, and transmit the audio data to the security management module;
[0011] A first storage unit, configured to cache the image frames to be displayed output by the first processing unit and transmit the image frames to be displayed to the security management module under the control of the first processing unit;
[0012] The first processing unit is further configured to control the first storage unit to transmit the image frames to be displayed to a video output processing unit when the security identification result is a first result, and transmit the audio data to an audio output processing unit;
[0013] A video output processing unit, connected to the first storage unit and the image display module, configured to perform a first processing on the image frames to be displayed and then output them to the image display module for display, where the first processing includes format conversion;
[0014] An audio output processing unit, connected to the first processing unit and the voice playback module, configured to perform a second processing on the audio data and then output it to the voice playback module for playback, where the second processing includes digital-to-analog conversion processing and amplification processing.
[0015] Optionally, the control module further includes a second processing unit, and the first processing unit is further configured to perform security processing on the audio data and control the second processing unit to perform security processing on the image frames to be processed when the security identification result is a second result.
[0016] Optionally, the first processing unit includes an expandable interface, and the security management module is connected to the first processing unit through the expandable interface.
[0017] Optionally, the security management module is a central processing unit, a graphics processing unit, a neural network processing unit, a system-on-chip, or a digital signal processor.
[0018] Optionally, the security management module includes at least one of an image security management unit, a text security management unit, and a voice security management unit;
[0019] The image security management unit is used to identify whether the image in the image frame to be displayed is secure content and generate a first identification result;
[0020] The text security management unit is used to identify whether the text in the image frame to be displayed is secure content and generate a second identification result;
[0021] The voice security management unit is used to identify whether the audio data is secure content and generate a third identification result;
[0022] The security identification result includes at least one of the first identification result, the second identification result, and the third identification result.
[0023] Optionally, the image security management unit is used to: perform intent recognition on the image frame to be displayed through machine learning or deep learning algorithms to generate a first keyword, compare the first keyword with a preset first keyword set, and determine whether the image in the image frame to be displayed is secure content according to the comparison result;
[0024] The text security management unit is used to: identify the text in the image frame to be displayed through an optical character recognition algorithm to obtain first text information, perform intent recognition on the first text information through a natural language processing algorithm to obtain a second keyword, compare the second keyword with a preset second keyword set, and determine whether the text in the image frame to be displayed is secure content according to the comparison result;
[0025] The voice security management unit is used to: convert the audio data into second text information through an automatic speech recognition algorithm, perform intent recognition on the second text information through a natural language processing algorithm to obtain a third keyword, compare the third keyword with a preset third keyword set, and determine whether the audio data is secure content according to the comparison result.
[0026] Optionally, the preset first keyword set, the preset second keyword set, and the preset third keyword set each include multiple keyword groups, each keyword group corresponds to a playback mode, and any one of the keyword groups includes multiple keywords;
[0027] The control module is further used to transmit the currently set playback mode by the user and the data to be played to the security management module together;
[0028] The security management module is further used to find the keyword group corresponding to the current playback mode from the preset first keyword set, the preset second keyword set, and the preset third keyword, and perform security identification on the data to be played by using the found keyword group.
[0029] Optionally, the security management module further includes:
[0030] A fusion unit, configured to perform weighted summation on at least two of the first recognition result, the second recognition result, and the third recognition result according to preset weights of the first recognition result, the second recognition result, and the third recognition result, and use the weighted summation result as the security recognition result.
[0031] Optionally, the security processing of the to-be-played data by the control module includes:
[0032] Determine a target area in the to-be-displayed image frame according to the first recognition result, and perform occlusion or replacement on the target area in the to-be-displayed image frame or the to-be-displayed image frame where the target area is located;
[0033] Determine target text in the to-be-displayed image frame according to the second recognition result, and perform occlusion or replacement on the target text in the to-be-displayed image frame;
[0034] Determine target audio data in the audio data according to the third recognition result, and delete or replace the target audio data.
[0035] Optionally, the control module further includes a transmission interface, connected to the first processing unit, for transmitting the to-be-played file to the first processing unit.
[0036] A second aspect of the present disclosure provides a playback control method, including the following steps:
[0037] Obtain a to-be-played file and decode the to-be-played file to obtain to-be-played data, where the to-be-played data includes to-be-displayed image frames and / or audio data;
[0038] Perform security recognition on the to-be-played data and generate a security recognition result;
[0039] When the security recognition result is a first result, play the to-be-played data;
[0040] When the security recognition result is a second result, perform security processing on the to-be-played data, and play the to-be-played data after the security processing.
[0041] The beneficial effects of the present disclosure are as follows:
[0042] After the playback terminal according to the embodiments of the present disclosure obtains the file to be played and decodes it to obtain the data to be played, first, the security management module performs security identification on the data to be played. After the security identification passes, it is played through the image display module and / or the voice playback module. If the security identification fails, the data to be played needs to be securely processed and then played. With such a setting, on the one hand, it is possible to perform real-time security identification on the playback content of the file to be played, intercept or replace the playback content before the picture is displayed on the display screen or the sound is played through the speaker, ensuring the security of the playback content. On the other hand, since the security identification process is automatic and is executed for all files to be played obtained by the playback terminal, the security monitoring efficiency is high, and it is possible to perform security identification on the files to be played in different playback scenarios, effectively avoiding the playback of illegal content caused by the illegal operation, operation error or intentional act of the playback personnel. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The following further describes in detail the specific embodiments of the present disclosure with reference to the drawings.
[0044] Figure 1 It is a block diagram of one of the structures of the playback terminal provided by the embodiments of the present disclosure;
[0045] Figure 2 It is a schematic diagram of one of the hardware structures of the playback terminal provided by the embodiments of the present disclosure;
[0046] Figure 3 It is a block diagram of one of the structures of the control module and the security management module provided by the embodiments of the present disclosure;
[0047] Figure 4 It is a schematic diagram of the pluggable connection between the security management module and the main control board provided by the embodiments of the present disclosure;
[0048] Figure 5 It is a schematic diagram of one of the hardware structures of an LCD display screen provided by the embodiments of the present disclosure;
[0049] Figure 6 It is a schematic diagram of the process of the playback terminal provided by the embodiments of the present disclosure for controlling the playback of audio and video files;
[0050] Figure 7 It is a flowchart of one of the embodiments of the playback control method provided by the embodiments of the present disclosure;
[0051] Figure 8 It is a flowchart of the playback control method provided by the embodiments of the present disclosure for performing security identification on images;
[0052] Figure 9 It is a flowchart of the playback control method provided by the embodiments of the present disclosure for performing security identification on text;
[0053] Figure 10 Flowchart for the secure recognition of speech by the playback control method provided by the embodiments of the present disclosure. Detailed implementation manners
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are only a part rather than all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0055] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The terms "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. Similarly, the terms such as "a", "an", or "the" do not denote a quantity limitation, but mean that there is at least one. The terms such as "comprising" or "including" mean that the elements or items appearing before the term cover the elements or items listed after the term and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0056] To better understand the technical solutions of the playback terminal and the playback control method of the present disclosure, the design concept of the present disclosure will be briefly introduced first.
[0057] The reasons for security incidents in the playback content of a display device include, but are not limited to, the following aspects: First, being attacked by the network, where hackers invade the playback system through technical means to tamper with or insert illegal content; second, external personnel using the screen mirroring function to play illegal content; third, the playback personnel play illegal content due to illegal operations, operational errors, or intentional actions. Therefore, how to effectively detect and prevent the security risks of playback content has become an urgent technical problem to be solved.
[0058] Currently, for the security risks of the content played by display devices, common solutions include manual monitoring, automatic monitoring, and video annotation. However, these solutions all have obvious deficiencies. Specifically, although manual monitoring can discover potential risks to a certain extent, due to its lack of timeliness, it often fails to discover and handle the playback of illegal content in a timely manner, resulting in a lag in the handling of accidents and an inability to effectively contain the expansion of the situation. Moreover, manual monitoring is inefficient and lacks flexibility, making it difficult to cope with large-scale and high-frequency playback requirements. The automatic monitoring solution is to add external monitoring devices, such as cameras and monitoring hosts, and use built-in software algorithms to monitor and judge the played content in real time. However, this solution also has many defects. First of all, the monitoring device can only collect and judge after the illegal content has been played, so there is a certain lag. Secondly, the playback device and the monitoring device are two independent systems and need to be connected wirelessly or wiredly, which not only increases the complexity of the system but also provides an opportunity for hackers to attack. For example, hackers can bypass the detection of the monitoring system by blocking the camera, shielding or changing the camera content, or even directly attacking the monitoring host. In addition, for deliberately internal personnel or short-term and non-repeating illegal content, the automatic monitoring solution often fails to effectively block and give early warnings. Although the video annotation solution can add specific watermarks to the video to prevent the content from being tampered with, this method is only applicable to the advertising scenario and cannot be applied to other types of played content, and the watermark technology itself also has the risk of being broken.
[0059] In summary, the existing security solutions for played content have deficiencies in terms of monitoring efficiency, timeliness, and application scenarios, and cannot meet the current high requirements for the security of played content.
[0060] To solve the above technical problems, the embodiments of the present disclosure provide a playback terminal and a playback control method. The following will describe the specific embodiments of the present disclosure in detail with reference to the accompanying drawings.
[0061] Please refer to Figure 1 , Figure 1 which is the structural block diagram of the playback terminal provided by the embodiments of the present disclosure. As Figure 1 shown, the playback terminal includes a control module 10, a security management module 20, an image display module 30, and a voice playback module 40. The control module 10 is connected to the security management module 20, the image display module 30, and the voice playback module 40.
[0062] The control module 10 is used to decode the file to be played to obtain the data to be played, transmit the data to be played to the security management module 20 and receive the security identification result fed back by the security management module 20, and when the security identification result is the first result, transmit the data to be played to the image display module 30 and / or the voice playback module 40 for playback. When the security identification result is the second result, perform security processing on the data to be played, and transmit the securely processed data to be played to the image display module 30 and / or the voice playback module 40 for playback. The data to be played includes the image frames to be displayed and / or audio data.
[0063] The security management module 20 is used to perform security identification on the data to be played and feed back the security identification result to the control module 10.
[0064] Optionally, the playback terminal in the embodiments of the present disclosure may be a display device such as an intelligent all-in-one machine (electronic whiteboard), an advertising screen, a computer monitor, a portable computer, a mobile phone, etc. Please refer to Figure 2 , Figure 2 is a schematic hardware structure diagram of an embodiment of a playback terminal. As Figure 2 shown, the basic components of the playback terminal include a display screen, a power supply board, a main control board, and a speaker. Among them, the display screen may be of types such as an LED display screen, an LCD display screen, an OLED display screen, etc., which are not limited here. Exemplarily, please refer to Figure 5 , Figure 5 is a schematic hardware structure diagram of an embodiment of an LCD display screen. As Figure 5 shown, the LCD display screen includes a backlight board, a TCON board, and an LCD screen. The backlight board provides light sources for the LCD screen, the TCON board provides recognizable timing control signals for the LCD screen, and the LCD screen drives the liquid crystal to rotate and transmit light to display image content. In addition, the TCON board may be integrated into the main control board, or the functions of the TCON board may be integrated into the main control board to form the technical effect of TCON Less; the function of the power supply board is to convert AC alternating current into DC direct current and supply power to the main control board and the display screen; the main control board is the control center of the playback terminal, which can output video signals and drive the display screen to display content, and output analog audio signals and drive the speaker to emit sound. Among them, the display screen is equivalent to Figure 1 the image display module 30 in Figure 1 , and the speaker is equivalent to Figure 1 the voice playback module 40 in
[0065] In the embodiments of the present disclosure, the security recognition result includes a first result and a second result. Among them, the first result is used to indicate that the playback content corresponding to the data to be played is secure content, and at this time, the security recognition passes; the second result is used to indicate that there are risks in the playback content corresponding to the data to be played, that is, there is content that is not suitable for playback, and at this time, the security recognition fails.
[0066] Among them, the playback content on the playback terminal usually includes two types: images and voices, and the images can also include text, such as subtitles. At this time, the security recognition result needs to include the security recognition results for different types of playback content, such as the security recognition result for images, the security recognition result for sounds, and the security recognition result for the text in the images.
[0067] Exemplarily, when the file to be played is an audio-visual file, at this time, the playback content includes video images and voices, and subtitles are usually included in the video images. At this time, for the video images, voices, and subtitles, security recognition needs to be performed respectively. Therefore, the security recognition result will include the recognition results for the video images, voices, and subtitles. According to the security recognition result, for the video images, voices, and subtitles, separate processing can also be performed. For example, when the images in the video images are secure content, the images are directly displayed on the image display module 30; when there is content in the video images that is not suitable for playback, the images are securely processed before playback; similarly, when the subtitles in the video images are secure content, the subtitles in the images are directly displayed on the image display module 30; when there is content in the video images that is not suitable for playback, the subtitles are securely processed before playback; when the voice is secure content, it is directly played in the voice playback module 40; when there is content in the voice that is not suitable for playback, the voice is securely processed before playback.
[0068] It can be understood that the file to be played can also be various types of files, such as an audio file, or a PPT file without voice, etc. However, when various types of files are played on the playback terminal, they will all be output in the form of images, sounds, or images and sounds. Therefore, in the embodiments of the present disclosure, the recognition of images, voices, and text can cover the security recognition of various types of files to be played.
[0069] Compared with the related art, after the playback terminal according to the embodiments of the present disclosure obtains the file to be played and decodes it to obtain the data to be played, first, the security management module is used to perform security identification on the data to be played. After the security identification is passed, the data is played through the image display module and / or the voice playback module. If the security identification fails, the data to be played needs to be processed securely, and only after the security processing will it be played. With such a setting, on the one hand, the playback content of the file to be played can be securely identified in real time, and the playback content can be intercepted or replaced before the picture is displayed on the display screen or the sound is played through the speaker, ensuring the security of the playback content. On the other hand, since the security identification process is automatic and is executed for all files to be played obtained by the playback terminal, the security monitoring efficiency is high, and the security identification of the files to be played can be realized in different playback scenarios, effectively avoiding the playback of illegal content caused by the illegal operation, operation error or intentional act of the playback personnel.
[0070] In a possible implementation manner, as Figure 3 shown, the control module 10 includes a first processing unit 110, a first storage unit 120, a video output processing unit 130, and an audio output processing unit 140.
[0071] The first processing unit 110 is configured to decode the file to be played and generate image data and / or audio data, divide the image data into image frames to be displayed and store them in the first storage unit 120, and transmit the audio data to the security management module 20.
[0072] Optionally, the image frame to be displayed is a picture format obtained by taking a screenshot of the image to be displayed. Exemplarily, when the file to be played is a video file, the image frame to be displayed is an image obtained by taking screenshots of each frame of the video image. Optionally, when the file to be played is an audio file or an audio-video file, the audio data obtained after decoding the audio file or the audio-video file is digital audio data.
[0073] The first storage unit 120 is configured to cache the image frames to be displayed output by the first processing unit 110 and transmit the image frames to be displayed to the security management module 20 under the control of the first processing unit 110.
[0074] The first processing unit 110 is further configured to control the first storage unit 120 to transmit the image frames to be displayed to the video output processing unit 130 and transmit the audio data to the audio output processing unit 140 when the security identification result is the first result.
[0075] A video output processing unit 130, connected to the first storage unit 120 and the image display module 30, is configured to perform a first processing on the image frame to be displayed and then output it to the image display module 30 for display. The first processing includes format conversion.
[0076] An audio output processing unit 140 is connected to the first processing unit 110 and the voice playback module 40, and is configured to perform a second processing on the audio data and then output it to the voice playback module 40 for playback. The second processing includes digital-to-analog conversion processing and amplification processing.
[0077] Optionally, the first processing unit 110 may be a central processing unit (CPU), a graphics processing unit (GPU), or a neural network processing unit (NPU). Among them, the CPU is a general-purpose processor, usually used to complete logical operations and fixed-point operations, and it has strong flexibility and programmability; although the GPU was originally designed for processing graphics and video rendering, it also has a large number of computing units and is suitable for performing large-scale parallel computing; the NPU is a processor specifically designed for accelerating neural network algorithms, with the characteristics of high performance and low energy consumption. It can be optimized for AI computing at the hardware level, significantly improving computing performance and energy efficiency. The NPU performs well in AI machine learning, especially suitable for scenarios that require a large amount of parallel computing, such as image recognition, speech recognition, and natural language processing. Exemplarily, in one embodiment, since the first processing unit 110 needs to perform a large number of logical operations, it is set as a central processing unit.
[0078] In the embodiments of the present disclosure, the first processing unit 110 decodes the file to be played and generates image data and / or audio data, including the following scenarios: when the file to be played is an audio file, the first processing unit 110 decodes the file to be played and obtains audio data; when the file to be played is a video file, a PPT or other file without audio, the first processing unit 110 decodes the file to be played and obtains image data; when the file to be played is an audio-visual file, the first processing unit 110 decodes the file to be played and obtains image data and audio data. Among them, for a PPT, each slide can be regarded as a "static frame", which contains visual elements such as text, pictures, and charts for conveying information to the audience; if the PPT contains animation effects, then these animation effects can be regarded as being realized in a series of "dynamic frames", and these dynamic frames are played at a certain time interval to present a coherent animation effect; and similar to video playback, the playback process of the PPT also shows each page or animation effect in a certain time sequence. Therefore, a PPT file can also be understood as being composed of multiple image frames. Similarly, for other types of files such as Word, when played, they are also displayed in the form of images, and each page can be understood as a frame of image. Therefore, the first processing unit 110 can decode and segment files of types such as PPT and Word to obtain multiple image frames to be displayed.
[0079] Among them, in the process of decoding the file to be played and generating image data and / or audio data, a decoder needs to be used. The decoder is used to decode the compressed audio-visual signal into the original audio-visual signal, and it can be a hardware decoder or a software decoder. When the decoder is a hardware decoder, it can be a dedicated decoding chip. When the decoder is a software decoder, it can be completed by a decoding program on the CPU. In the embodiments of the present disclosure, the first processing unit 110 can complete the decoding process of the file to be played. It can be understood that in other embodiments, an independent codec can also be set, and the independent codec decodes the file to be played, and the obtained image data is segmented into image frames, and each frame of image is stored in the first storage unit 120.
[0080] Optionally, the first storage unit 120 is used for data storage and temporary storage, including but not limited to types such as Double Data Rate Synchronous Dynamic Random Access Memory (DDR for short), Embedded Multi Media Card (eMMC for short), and Universal Flash Storage (UFS for short).
[0081] Optionally, the video output processing unit 130 is a VOP (Video Output Processor), which is mainly used to perform a first processing on the image frame to be displayed and then output it to the display screen. Exemplarily, the first processing includes format conversion, such as converting the received image frame to be displayed into a format suitable for display on the display screen, including but not limited to adjusting parameters such as resolution, color depth, refresh rate, etc., to ensure the compatibility between the image frame to be displayed and the display screen. The VOP outputs the image frame to be displayed after the first processing to a video output interface, such as HDMI, DVI interface, etc., which are connected to a timing controller (TCON) board or other driver chips, and finally drives the display screen to display the image. In addition, in some cases, the VOP is also used to ensure the synchronous output of video signals and audio signals to provide a smooth and consistent audio-visual experience.
[0082] Optionally, the audio output processing unit 140 includes a digital-to-analog conversion circuit and an amplification circuit. The digital-to-analog conversion circuit is used to convert the received audio data into analog audio data and output it to the amplification circuit, and the amplification circuit is used to amplify the analog audio data and then output it to the speaker to drive the speaker to play sound. Specifically, the first processing unit 110 outputs the audio data after security identification through the audio output interface, converts it into analog audio data through DAC, and then performs power amplification through the amplification circuit, and finally drives the speaker to play sound.
[0083] Optionally, as Figure 3 shown, the control module 10 further includes a transmission interface 150, which is connected to the first processing unit 110 and is used to transmit the file to be played to the first processing unit 110; the transmission interface includes one or more of a USB (Universal Serial Bus) interface, an Ethernet interface, a PCIe (PCI Express) interface, an SPI (Serial Peripheral Interface) interface, an HDMI (High-Definition Multimedia Interface) interface, a VBO (V-by-One) interface, an eDP (Embedded DisplayPort) interface, a MIPI interface (Mobile Industry Processor Interface), a DP interface (DisplayPort), an IIS interface (Integrated Audio Interface), a MIC ADC (Microphone Analog-to-Digital Converter) interface.
[0084] Exemplarily, the transmission interface 150 includes an input / output interface 1510 and an audio / video input interface 1520. The input / output interface 1510 includes, but is not limited to, interfaces such as USB, Ethernet, PCIe, SPI, etc. The input / output interface 1510 provides the data exchange capability between the playback terminal and external devices, and is used to complete the input or output of data between the playback terminal and external devices. The audio / video input interface 1520 is responsible for the input and output of audio / video, connects the playback terminal and an external audio / video source, and includes, but is not limited to, interface types such as HDMI, VBO, eDP, MIPI, DP, IIS, MIC ADC, etc., and is used to complete the input or output of audio / video data.
[0085] In the embodiments of the present disclosure, the playback terminal has various types of transmission interfaces 150. Therefore, the file to be played can have various sources, such as locally stored videos (disks, USB drives, etc.), network video streams (Internet, local NAS, NVR, etc.), local video information (HDMI, DP, VGA, etc. wired or wireless screen mirroring such as Miracast, DLNA, Airplay, etc.), locally acquired information (Camera, MIC, etc.). At this time, the secure identification of the file to be played in various playback scenarios can be realized, such as the secure identification of network videos and network audio, the audio input from the microphone in a live broadcast or meeting, or the secure identification of remote audio / video content.
[0086] In a possible implementation manner, the control module 10 further includes a second processing unit 160. The first processing unit 110 is further configured to perform security processing on the audio data when the security identification result is the second result, and control the second processing unit 160 to perform security processing on the image frame to be processed.
[0087] In the embodiments of the present disclosure, the first processing unit 110 is a CPU, and the second processing unit 160 can be a GPU or an NPU. The first processing unit 110 and the second processing unit 160 cooperate with each other to realize the secure management of the playback content. Exemplarily, in the security identification result: when there is insecure content in the audio data, the first processing unit 110 is used to perform security processing on the audio data; when there is insecure content in the image frame to be displayed, the second processing unit is used to perform security processing on the image frame to be displayed. With such a setting, the parallel computing capabilities of the GPU or NPU can be effectively utilized, the processing speed and efficiency can be improved, and according to the task type (audio processing or image processing), the task can be assigned to the most suitable processing unit, which can more effectively utilize the system resources. This way of division of labor and cooperation avoids the situation of overloading a single processing unit and improves the stability and response speed of the entire system.
[0088] In a possible implementation, the main control board is provided with an expandable interface 170, and the access of the security management module 20 is realized through the expandable interface 170. Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of the pluggable connection between the security management module 20 and the main control board provided by the embodiment of the present disclosure.
[0089] Optionally, multiple expansion interfaces can be reserved on the main control board, including but not limited to M.2, MiniPCIe, USB, SDIO interfaces, etc. These interfaces can provide users with rich choices to adapt to different application scenarios and computing power requirements. Exemplarily, an M.2 expansion interface is reserved in the embodiment of the present disclosure. The M.2 interface has the advantages of high-speed transmission, stability and reliability, easy plugging and unplugging, etc., and is suitable for connecting the high-performance security management module 20. Optionally, the security management module 20 is a processor or controller with computing power, including but not limited to: central processing unit, graphics processing unit, neural network processing unit, system-on-chip (SoC) or digital signal processor (DSP), etc. Users can select the most suitable security management module 20 according to actual needs and insert it into the main control board through the reserved M.2 expansion interface (or other expansion interfaces).
[0090] During specific implementation, the first processing unit 110 on the main control board includes an expandable interface 170, and the security management module 20 is connected to the first processing unit 110 through the expandable interface 170.
[0091] The embodiment of the present disclosure realizes flexible expansion and change of computing power by reserving the expandable interface 170 and adopting the pluggable security management module 20. This solution not only improves the flexibility and scalability of the playback terminal, but also provides users with a rich selection space. Users can conveniently upgrade or replace the security management module to adapt to new application scenarios and computing power requirements.
[0092] In a possible implementation, as Figure 3 shown, the control module 10 further includes a power management module 180, and the power management module 180 is used to provide power for each module in the main control board and output a stable supply voltage.
[0093] In a possible implementation, as Figure 3 shown, the control module 10 may further include a third processing unit 190, where one of the third processing unit 190 and the second processing unit 160 is a GPU and the other is an NPU.
[0094] In a possible implementation, the security management module 20 includes at least one of an image security management unit 210, a text security management unit 220, and a voice security management unit 230. The image security management unit 210 is configured to identify whether the image in the to-be-displayed image frame is secure content and generate a first identification result; the text security management unit 220 is configured to identify whether the text in the to-be-displayed image frame is secure content and generate a second identification result; the voice security management unit 230 is configured to identify whether the audio data is secure content and generate a third identification result; the security identification result includes at least one of the first identification result, the second identification result, and the third identification result.
[0095] Exemplarily, as Figure 3 shown, the security management module 20 includes an image security management unit 210, a text security management unit 220, and a voice security management unit 230. At this time, security identification of video, text, and audio content can be realized. The text here refers to the text in the image when it is displayed.
[0096] In a possible implementation, the image security management unit 210 is configured to: perform intent recognition on the to-be-displayed image frame through a machine learning or deep learning algorithm to generate a first keyword, compare the first keyword with a preset first keyword set, and determine whether the image in the to-be-displayed image frame is secure content according to the comparison result.
[0097] Taking the to-be-played file as a video file as an example, after the first processing unit 110 decodes and performs frame-by-frame segmentation on the to-be-played file, multiple to-be-displayed image frames included in the video file can be obtained and cached in the first storage unit 120. When the image security management unit 210 performs intent recognition on the to-be-displayed image frame, it performs intent recognition on multiple to-be-displayed image frames cached in the first storage unit 120.
[0098] Optionally, machine learning (SVM) or deep learning algorithms can be used to perform intent recognition on the to-be-displayed image frame. The deep learning algorithm can be, for example, a convolutional neural network (CNN), RCNN, FastText, etc. In specific implementation, first, it is necessary to use sample images (such as a sample video including multiple consecutive frames of images) to train an image intent recognition model, and then use the trained image intent recognition model to perform intent recognition on the to-be-displayed image frame to generate a label representing the intent of the to-be-displayed image frame, and this label is the first keyword.
[0099] Exemplarily, the process of training an image intention recognition model using sample videos includes but is not limited to the following: (1) Manually annotate the videos and assign one or more intention tags to each video; (2) Feature extraction: Split the sample videos into a series of video frames, perform image processing on each video frame, and extract visual features such as color, shape, and texture. Specifically, a deep learning model (such as a convolutional neural network CNN) can be used for feature extraction to capture more complex image features; (3) Temporal feature extraction: Consider the temporal relationship between video frames and extract temporal features such as motion trajectories and speed changes. Specifically, models such as recurrent neural networks (RNN) or long short-term memory networks (LSTM) can be used to process temporal data. (3) Model training: Select a suitable machine learning model or deep learning model according to the task requirements and data characteristics, and use the annotated sample videos to train the model. During the training process, continuously adjust the model parameters to minimize the loss function. In the embodiments of the present disclosure, the image security management unit 210 processes each frame of the image, so it can identify the bad information in the entire file to be played.
[0100] It can be understood that the embodiments of the present disclosure utilize the technology of intention recognition to perform intention recognition on the image frames to be displayed, but do not improve the technology of intention recognition itself. The specific algorithms for intention recognition can adopt mature technologies in related technologies.
[0101] In a possible implementation manner, the text security management unit 220 is configured to: identify the text in the image frame to be displayed through an optical character recognition algorithm (OCR) to obtain first text information, perform intention recognition on the first text information through a natural language processing algorithm (NLP) to obtain second keywords, compare the second keywords with a preset second keyword set, and determine whether the text in the image frame to be displayed is secure content according to the comparison result.
[0102] In the embodiments of the present disclosure, for some specific occasions, it is required that specific keywords or certain specific content cannot be displayed on the display screen. At this time, the OCR technology can be used to perform character recognition on the image frames to be displayed to obtain the text information contained in each image frame to be displayed, that is, the first text information, and then the NLP technology is used to perform intention recognition on the first text information to obtain second keywords, where the second keywords can represent the intention or theme expressed by the first text information.
[0103] Among them, intent recognition is a key task in NLP semantic analysis, which can be used to judge the intent or theme expressed in the text. The methods of intent recognition can be divided into three categories: rule-based, template-based, and deep learning-based. The embodiments of the present disclosure do not improve the specific implementation algorithm of NLP intent recognition, but use the existing NLP intent recognition algorithm to perform intent recognition on the text in the image to be displayed (i.e., the first text information), and then judge whether there is illegal content in the text of the image to be displayed. When it is determined that the text in the image to be displayed is not safe content, it cannot be directly displayed.
[0104] In other embodiments, the first text information can also be securely identified through keyword comparison technology. Exemplarily, it can be retrieved whether there are keywords in the preset second keyword set in the first text information. If not, it is determined that the text in the image frame to be displayed is safe content. Otherwise, it is determined that there are sensitive words in the text of the image frame to be displayed and security processing is required.
[0105] In a possible implementation manner, the voice security management unit 230 is configured to: convert the audio data into second text information through an automatic speech recognition algorithm (Automatic Speech Recognition, abbreviated as ASR), perform intent recognition on the second text information through a natural language processing algorithm (NLP) to obtain a third keyword, compare the third keyword with a preset third keyword set, and judge whether the audio data is safe content according to the comparison result.
[0106] In the embodiments of the present disclosure, NLP is used to perform intent recognition on the second text information to obtain a third keyword, where the third keyword can represent the intent or theme expressed by the second text information. The embodiments of the present disclosure do not improve the specific implementation algorithm of NLP intent recognition, but use the existing NLP intent recognition algorithm to perform intent recognition on the text information corresponding to the audio data (i.e., the second text information), and then judge whether there is illegal content in the audio data. If it is determined that the audio data is not safe content, it cannot be directly played.
[0107] Performing intent recognition on the audio content to be played can analyze the potential intent in the audio content, thereby effectively identifying bad or illegal information, ensuring that the audio content to be played complies with laws, regulations and social ethics, avoiding the spread of harmful information, and protecting the public from the infringement of bad information. In addition, intent recognition can also quickly identify the key information and potential risks in the audio content, realize automated and intelligent risk recognition, and improve the security of the played content.
[0108] In the embodiments of the present disclosure, the image security management unit 210, the text security management unit 220, and the voice security management unit 230 all understand the playback content in the file to be played, such as images, texts, or audios, through the intention recognition method, and confirm whether the playback content is legal and conforms to the filtered content settings, so as to achieve the effect of secure playback and prevent illegal content from being played on the playback terminal by illegal personnel.
[0109] In a possible implementation manner, the preset first keyword set, the preset second keyword set, and the preset third keyword set may be a keyword blacklist or a keyword whitelist.
[0110] The keyword blacklist is a database containing illegal or sensitive keywords, and these keywords can be defined according to the requirements of laws, ethics, or specific application scenarios. When the first keyword is a certain keyword in the keyword blacklist, it indicates that the illegal content set in the keyword blacklist is included in the image frame to be displayed at this time. At this time, the image frame to be displayed cannot be directly played and needs to be securely processed.
[0111] The keyword whitelist includes keywords representing secure content, and these keywords are usually related to corporate promotional videos, advertisements for specific content, etc. If the first keyword only contains the keywords set in the keyword whitelist, it indicates that the content in the image frame to be displayed is secure at this time and can be directly played; if the first keyword contains keywords other than those in the keyword whitelist, it indicates that the illegal content is included in the image frame to be displayed at this time. At this time, the image frame to be displayed cannot be directly played and needs to be securely processed.
[0112] It can be understood that the preset first keyword set, the preset second keyword set, and the preset third keyword set can be personalized according to the application scenario of the playback terminal. The keywords in the preset first keyword set, the preset second keyword set, and the preset third keyword set can be the same or different, and the embodiments of the present disclosure do not limit them.
[0113] In a possible implementation manner, the preset first keyword set, the preset second keyword set, and the preset third keyword set all include multiple keyword groups, each keyword group corresponds to a playback mode, and any one of the keyword groups includes multiple keywords; the control module is further configured to transmit the currently set playback mode by the user and the data to be played to the security management module together; the security management module is further configured to find the keyword group corresponding to the currently set playback mode from the preset first keyword set, the preset second keyword set, and the preset third keyword set, and perform security identification on the data to be played by using the found keyword group.
[0114] In the embodiments of the present disclosure, multiple playback modes can be set for the playback terminal, such as children's playback mode, adolescent playback mode, adult playback mode, etc. In different playback modes, the safe content allowed to be played by the playback terminal is also different. Therefore, it is necessary to preset keywords for different playback modes. Specifically, when implementing, for different playback modes, keyword groups for the corresponding playback modes are set. When the playback terminal is in different playback modes, the keyword groups are used to perform safety identification on the playback content. Similarly, the keyword groups can be white lists or black lists of keywords.
[0115] In a possible implementation manner, as Figure 3 shown, the security management module 20 further includes a fusion unit 240. The fusion unit 240 is configured to perform weighted summation on at least two of the first recognition result, the second recognition result, and the third recognition result according to preset weights of the first recognition result, the second recognition result, and the third recognition result, and use the weighted summation result as the security recognition result.
[0116] Optionally, when the file to be played covers both voice and image at the same time (for example, the file to be played is an audio-video file, or the file to be played includes a PPT and an audio file input by a microphone), if it is necessary to identify whether the playback content is safe content, on the basis of discriminating the voice and the image respectively, a fusion discrimination algorithm for the voice and the image can be added, that is, by comprehensively fusing the recognition results of different types of playback content for judgment. Since the final judgment result comprehensively considers the security of the image and voice content, a more accurate playback decision can be made.
[0117] Exemplarily, assuming that the fusion unit 240 is used to fuse the first recognition result and the third recognition result, the weight of the image content security recognition result can be set as a, and the weight of the voice content security recognition result can be set as b. Assuming that the image content security recognition result is δ1 and the voice content security recognition result is δ2, the fused security recognition result is expressed as: (δ1*a) / (a + b)+(δ2*b) / (a + b). Then, the security recognition result is compared with a preset security threshold (for example, 0.8 or other values determined according to actual requirements). If the fused security recognition result is greater than or equal to this security threshold, it is determined that the playback content is safe, and the image and voice can be played normally. If the fused security recognition result is less than this security threshold, the discrimination fails, and the image and voice cannot be directly played and need to be processed securely.
[0118] In the embodiments of the present disclosure, by introducing weights a and b, the importance of images and voices in analysis can be flexibly adjusted, so as to more accurately reflect the actual security requirements. Exemplarily, assuming that the security recognition accuracy of images is higher while that of voices is weaker, the preset weight a of the first recognition result can be set to a larger value, and the preset weight b of the second recognition result can be set to a value smaller than a. At the same time, since information of multiple modalities (images and voices) is fused, even if the information of one modality is interfered with or inaccurate, the information of the other modality can still make up for it to a certain extent, thereby improving the overall recognition accuracy and avoiding misjudgment. It should be noted that the values of weights a and b should be reasonably set according to the actual application scenarios and requirements. At the same time, the setting of the security threshold also needs to be adjusted according to historical data and actual test results to ensure the accuracy and reliability of the security recognition of the played content.
[0119] It can be understood that in other embodiments, any two or three of the first recognition result, the second recognition result, and the third recognition result can also be fused and judged.
[0120] In a possible implementation manner, the security processing of the data to be played by the control module includes: determining the target area in the image frame to be displayed according to the first recognition result, and performing occlusion or replacement on the target area in the image frame to be displayed or the image frame to be displayed where the target area is located; determining the target text in the image frame to be displayed according to the second recognition result, and performing occlusion or replacement on the target text in the image frame to be displayed; determining the target audio data in the audio data according to the third recognition result, and deleting or replacing the target audio data.
[0121] In the embodiments of the present disclosure, when the security recognition result of the image in the image frame to be displayed fails, the image frame to be displayed cannot be directly played. It is necessary to stop playing or perform security processing before playing. In addition, the control module can also report illegal content. The target area is the area where the security recognition fails in the image frame to be displayed. For example, there is illegal content in area A of the entire frame image. At this time, area A is the target area. The security processing of the image frame to be displayed can include but is not limited to the following: performing occlusion or replacement on the local area where the illegal content is located or the entire frame image where the illegal content is located. The occlusion includes, for example, blackening, mosaicing, etc. The replacement can be replaced with a blank image or a preset security image.
[0122] In the embodiments of the present disclosure, when the security recognition result of the text in the image frame to be displayed fails, the image frame to be displayed cannot be directly played and needs to be securely processed before playing. In addition, the control module can also report illegal content. Herein, the target text is the character in the image to be played that fails the security recognition. At this time, the security processing of the image frame to be displayed may include, but is not limited to, the following: obscuring or replacing the target text, where obscuring includes, for example, blackening, blurring, etc., and replacement can be replacing with an empty character or a specific character.
[0123] In the embodiments of the present disclosure, when the security recognition result of the audio data fails, the audio data cannot be directly played and needs to stop playing or be securely processed before playing. In addition, the control module can also report illegal content. Herein, the target audio data is the part of the audio data that fails the security recognition. At this time, the security processing of the audio data may include, but is not limited to, the following: eliminating the target audio data, replacing it with a prompt tone, or directly filling it with a prompt tone, etc.
[0124] It can be understood that during the process of intent recognition of the image frame to be displayed or audio data, features need to be extracted and then the features are recognized. During the process of feature extraction, the location information of the features will be located. Therefore, in addition to carrying information indicating whether the recognition passes, the corresponding recognition result will also carry the specific location of the playing content that fails the security recognition, that is, the location where the above-mentioned target area, target text, and target audio data are located.
[0125] Exemplarily, the process of intent recognition of an image includes: (1) Feature extraction and analysis: Using computer vision technology, features of the image are extracted, such as face recognition. During the feature extraction process, location technologies such as bounding boxes are used to determine the locations of each feature. Then, through a deep learning model, the extracted features are used for intent recognition to identify sensitive features in the image, and the location of the sensitive feature is the target area.
[0126] Exemplarily, the process of intent recognition of a digital audio signal includes: converting the digital audio signal into a format suitable for processing by a machine learning model, such as text information; using natural language processing to perform intent recognition on the text information to determine whether there is insecure content. After identifying the insecure content, a timestamp technology is used to mark the time point when the insecure content appears, and through algorithm calculation, the specific location of the insecure content in the audio (such as the start time and end time) is determined.
[0127] The following combines Figure 6 to describe the working process of the playback terminal provided in the embodiments of the present disclosure, where Figure 6 taking the example of inputting a file to be played through an audio-video input interface and the file to be played being an audio-video file, as Figure 6As shown, after the file to be played is input from the audio-video input interface, it enters the audio-video decoder. After being decoded by the audio-video decoder, the processable video image data and audio data are obtained. The video image data is segmented into image frames (i.e., the images to be displayed), and each frame of the image is stored in the Framebuffer (i.e., the first storage unit). In the related art, the image frames in the Framebuffer are directly sent to the video output processing unit VOP and finally output to the display screen. However, in the embodiments of the present disclosure, the image frames in the Framebuffer are first sent by the control module to the security management module to complete image security recognition. Only after the security recognition passes will the image frames be output to the video output processing unit and output to the TCON board through the video output interface, and the display screen will be driven to display the content. Similarly, the audio data is transmitted in real time to the security management module for voice security recognition. Only after the security recognition passes will the audio data be output via the audio output interface and then enter the audio output processing unit. After being processed by the DAC, the audio data becomes analog audio data. After being amplified by the power amplifier, the speaker is driven to play the sound.
[0128] Based on the same inventive concept, the second aspect of the present disclosure provides a playback control method, which is applied to the playback terminal as described above. Please refer to Figure 7 , Figure 7 which is a schematic flowchart of the playback control method of the embodiments of the present disclosure. As Figure 7 shown, it includes the following steps:
[0129] S101, obtain the file to be played and decode the file to be played to obtain the data to be played, where the data to be played includes the images to be displayed and / or audio data;
[0130] S102, perform security recognition on the data to be played and generate a security recognition result. When the security recognition result is the first result, execute step S103. When the security recognition result is the second result, execute step S104.
[0131] S103, play the data to be played, that is, play the playback content corresponding to the data to be played;
[0132] S104, perform security processing on the data to be played, and play the data to be played after security processing, that is, play the playback content corresponding to the data to be played after security processing.
[0133] Among them, in step S101, the control module 10 obtains the file to be played, decodes the file to be played to obtain the data to be played, and transmits the data to be played to the security management module 20 for security identification. In step S102, the security management module 20 performs security identification on the data to be played and feeds back the security identification result to the control module 10. In steps S103 and S104, the control module 10 receives the security identification result fed back by the security management module 20. When the security identification result is the first result, the data to be played is transmitted to the image display module 30 and / or the voice playback module 40 for playback. When the security identification result is the second result, the data to be played is subjected to security processing, and the data to be played after security processing is transmitted to the image display module 30 and / or the voice playback module 40 for playback.
[0134] In a possible implementation manner, in step S102, the steps for the security management module 20 to perform security identification on the data to be played include at least one of the following steps (1) to (3):
[0135] (1) Identify whether the image in the image frame to be displayed is secure content and generate a first identification result;
[0136] (2) Identify whether the text in the image frame to be displayed is secure content and generate a second identification result;
[0137] (3) Identify whether the audio data is secure content and generate a third identification result;
[0138] At this time, the security identification result includes at least one of the first identification result, the second identification result, and the third identification result.
[0139] In a possible implementation manner, when the data to be played includes an image frame to be displayed, as Figure 8 shown, the playback control method includes the following steps:
[0140] Step S201, the first processing unit obtains the file to be played and decodes the file to be played to obtain image data.
[0141] Step S202, the first processing unit divides the image data into image frames to be displayed and stores them in the first storage unit for caching.
[0142] Step S203, the first processing unit controls the first storage unit to output the cached image frames to be displayed to the security management module.
[0143] Step S204, the security management module performs intent recognition on the image frames to be displayed through machine learning or deep learning algorithms and generates a first keyword.
[0144] Step S205: Compare the first keyword with a preset first keyword set, determine whether the image in the to-be-displayed image frame is safe content according to the comparison result, and feedback a first recognition result to the control module. If the determination result is yes, the first recognition result includes a first result at this time; if the determination result is no, the first recognition result includes a second result at this time.
[0145] Step S206: When the security recognition result is the first result, the control module controls the security management module to transmit the to-be-displayed image frame to the image display module for normal playback.
[0146] Step S207: When the security recognition result is the second result, the control module reports illegal content and executes Step S208.
[0147] Step S208: Control the image display module to stop playing or play the to-be-displayed image frame after security processing, where the security processing operations include: the control module determines a target area in the to-be-displayed image frame according to the first recognition result, and performs occlusion or replacement on the target area in the to-be-displayed image frame or the to-be-displayed image frame where the target area is located.
[0148] In a possible implementation, when the data to be played includes a to-be-displayed image frame, as Figure 9 shown, the playback control method further includes the following steps:
[0149] Step S301: The first processing unit obtains the file to be played and decodes the file to be played to obtain image data.
[0150] Step S302: The first processing unit divides the image data into to-be-displayed image frames and stores them in the first storage unit for caching.
[0151] Step S303: The first processing unit controls the first storage unit to output the cached to-be-displayed image frames to the security management module.
[0152] Step S304: The security management module identifies the text in the to-be-displayed image frame through an optical character recognition algorithm to obtain first text information.
[0153] Step S305: The security management module performs intent recognition on the first text information through a natural language processing algorithm to obtain a second keyword.
[0154] Step S306: Compare the second keyword with a preset second keyword set, determine whether the text in the to-be-displayed image frame is safe content according to the comparison result, and feedback a second recognition result to the control module. If the determination result is yes, the second recognition result includes a first result at this time; if the determination result is no, the second recognition result includes a second result at this time.
[0155] Step S307, when the security recognition result is the first result, the control module controls the security management module to transmit the to-be-displayed image frame to the image display module for normal playback.
[0156] Step S308, when the security recognition result is the second result, the control module reports the illegal content and executes Step S309.
[0157] Step S309, control the image display module to stop playing or play the to-be-displayed image frame after security processing, where the security processing operations include: the control module determines the target text in the to-be-displayed image frame according to the second recognition result, and performs occlusion or replacement on the target text in the to-be-displayed image frame.
[0158] In a possible implementation manner, when the data to be played includes audio data, as Figure 10 shown, the playback control method further includes the following steps:
[0159] Step S401, the first processing unit obtains the file to be played and decodes the file to be played to obtain audio data.
[0160] Step S402, the first processing unit outputs the audio data to the security management module.
[0161] Step S403, the security management module converts the audio data into second text information through an automatic speech recognition algorithm.
[0162] Step S404, the security management module performs intent recognition on the second text information through a natural language processing algorithm to obtain a third keyword.
[0163] Step S405, compare the third keyword with a preset third keyword set, judge whether the audio data is secure content according to the comparison result, and feedback the third recognition result to the control module. If the judgment result is yes, the third recognition result includes the first result at this time. If the judgment result is no, the third recognition result includes the second result at this time.
[0164] Step S406, when the security recognition result is the first result, the control module transmits the audio data to the voice playback module for normal playback.
[0165] Step S407, when the security recognition result is the second result, the control module reports the illegal content and executes Step S408.
[0166] Step S408, control the voice playback module to stop playing or play the audio data after security processing, where the security processing operations include: the control module determines the target audio data in the audio data according to the third recognition result, and deletes or replaces the target audio data.
[0167] In a possible implementation, the preset first keyword set, the preset second keyword set, and the preset third keyword set each include a plurality of keyword groups, each keyword group corresponding to a playback mode, and any one of the keyword groups including a plurality of keywords;
[0168] The playback control method according to the embodiment of the present disclosure further includes: the control module transmits the currently set playback mode by the user and the data to be played to the security management module together; the security management module is further configured to search for a keyword group corresponding to the currently set playback mode from the preset first keyword set, the preset second keyword set, and the preset third keyword, and perform security identification on the data to be played by using the found keyword group.
[0169] In a possible implementation, the playback control method further includes: the security management module performs weighted summation on at least two of the first identification result, the second identification result, and the third identification result according to preset weights of the first identification result, the second identification result, and the third identification result, and uses the weighted summation result as the security identification result.
[0170] Based on the same inventive concept, a third aspect of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the playback control method described above are implemented.
[0171] In a specific implementation process, the computer storage medium may include: various storage media capable of storing program codes such as a universal serial bus flash drive (USB, Universal Serial Bus Flash Drive), a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc.
[0172] Based on the same inventive concept, a fourth aspect of the present disclosure provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the playback control method described above are implemented. Since the principle of solving problems by the above computer program is similar to the principle of the playback control method, the implementation of the above computer program can refer to the implementation of the playback control method, and the repeated parts will not be described again.
[0173] A computer program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0174] Obviously, the above embodiments of the present disclosure are merely examples for clearly illustrating the present disclosure, rather than limitations on the implementation manners of the present disclosure. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is impossible to enumerate all the implementation manners here. Any obvious changes or variations derived from the technical solutions of the present disclosure still fall within the protection scope of the present disclosure.
Claims
1. A playback terminal, characterized in that: It includes a control module, a security management module, an image display module and a voice playback module, wherein the control module is connected to the security management module, the image display module and the voice playback module; The control module is used to decode the to-be-played file to obtain the to-be-played data, transmit the to-be-played data to the security management module and receive the security identification result fed back by the security management module, and transmit the to-be-played data to the image display module and / or the voice playback module for playback when the security identification result is a first result, perform security processing on the to-be-played data when the security identification result is a second result, and transmit the to-be-played data after security processing to the image display module and / or the voice playback module for playback, wherein the to-be-played data includes image frames to be displayed and / or audio data; The security management module is used to perform security identification on the data to be played and feed back the security identification result to the control module.
2. The playback terminal according to claim 1, characterized in that: The control module comprises: a first processing unit, configured to decode the file to be played and generate image data and / or audio data, divide the image data into image frames to be displayed and store them in a first storage unit, and transmit the audio data to the security management module; a first storage unit, configured to cache the image frames to be displayed output by the first processing unit and transmit the image frames to be displayed to the security management module under the control of the first processing unit; The first processing unit is further used to control the first storage unit to transmit the image frame to be displayed to the video output processing unit, and to transmit the audio data to the audio output processing unit when the security identification result is the first result; a video output processing unit connected to the first storage unit and the image display module, and configured to perform a first processing on the image frame to be displayed and then output it to the image display module for display, wherein the first processing includes format conversion; The audio output processing unit is connected to the first processing unit and the voice playing module, and is used to perform a second processing on the audio data and then output it to the voice playing module for playing, wherein the second processing includes a digital-to-analog conversion process and an amplification process.
3. The playback terminal according to claim 2, characterized in that: The control module also includes a second processing unit, and the first processing unit is further used to perform security processing on the audio data when the security identification result is a second result, and control the second processing unit to perform security processing on the image frame to be processed.
4. The playback terminal according to claim 2, characterized in that: The first processing unit includes an extensible interface, and the security management module is connected to the first processing unit via the extensible interface.
5. The playback terminal according to claim 4, characterized in that: The security management module is a central processing unit, a graphics processing unit, a neural network processing unit, a system-on-chip or a digital signal processor.
6. The playback terminal according to claim 1, characterized in that: The security management module includes at least one of an image security management unit, a text security management unit, and a voice security management unit; The image security management unit is used to identify whether the image in the image frame to be displayed is a security content and generate a first identification result; The text security management unit is used to identify whether the text in the image frame to be displayed is a security content and generate a second recognition result; The voice security management unit is used to identify whether the audio data is security content and generate a third identification result; The security identification result includes at least one of the first identification result, the second identification result, and the third identification result.
7. The playback terminal according to claim 6, characterized in that: The image security management unit is used to: perform intent recognition on the image frame to be displayed through a machine learning or deep learning algorithm and generate a first keyword, compare the first keyword with a preset first keyword set, and determine whether the image in the image frame to be displayed is a safe content according to the comparison result; The text security management unit is used to: identify the text in the image frame to be displayed by an optical character recognition algorithm to obtain first text information, identify the intent of the first text information by a natural language processing algorithm to obtain a second keyword, compare the second keyword with a preset second keyword set, and determine whether the text in the image frame to be displayed is a safe content according to the comparison result; The voice security management unit is used to: convert the audio data into second text information through an automatic speech recognition algorithm, perform intent recognition on the second text information through a natural language processing algorithm to obtain a third keyword, compare the third keyword with a preset third keyword set, and determine whether the audio data is safe content based on the comparison result.
8. The playback terminal according to claim 7, characterized in that: The preset first keyword set, the preset second keyword set and the preset third keyword set all include multiple keyword groups, each keyword group corresponds to a play mode, and any of the keyword groups includes multiple keywords; The control module is also used to transmit the current playback mode set by the user together with the data to be played to the security management module; The security management module is also used to search for a keyword group corresponding to the current playback mode from the preset first keyword set, the preset second keyword set and the preset third keyword set, and use the found keyword group to perform security identification on the data to be played.
9. The playback terminal according to claim 6, characterized in that: The security management module also includes: A fusion unit is used to perform weighted summation on at least two of the first recognition result, the second recognition result and the third recognition result according to preset weights of the first recognition result, the second recognition result and the third recognition result, and use the weighted summation result as the security recognition result.
10. The playback terminal according to claim 7, characterized in that: The control module performs security processing on the data to be played, including: Determine the target area in the image frame to be displayed according to the first recognition result, and block or replace the target area in the image frame to be displayed or the image frame to be displayed where the target area is located; Determine the target text in the image frame to be displayed according to the second recognition result, and cover or replace the target text in the image frame to be displayed; Determine target audio data in the audio data according to the third recognition result, and delete or replace the target audio data.
11. The playback terminal according to claim 2, characterized in that: The control module further includes a transmission interface, which is connected to the first processing unit and is used to transmit the file to be played to the first processing unit.
12. A playback control method, characterized in that: The following steps are involved: Acquire a file to be played and decode the file to be played to obtain data to be played, wherein the data to be played includes image frames to be displayed and / or audio data; Performing security identification on the data to be played and generating a security identification result; When the security identification result is the first result, playing the data to be played; When the security identification result is the second result, the data to be played is security processed, and the data to be played after the security processing is played.
Citation Information
Cited By
Display equipment safety management system
CN120811677A