Real-time threat sensing method and device for network media equipment and electronic equipment
By comparing the characteristic sequence of the played content and the current play content in real time in the network media device, the problem of vulnerability to attack and tampering with content in the network media device is solved, and high-accurate threat awareness and timely response are achieved to ensure content security.
Patent Information
- Application Number
- CN202411921107.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-16
AI Technical Summary
In public places, online media devices have become the key target of criminals' cyber attacks due to the large number of display screens and wide audiences. This may cause the content to be played or tampered with, spread false information or malicious propaganda, causing misleading and social panic.
A real-time threat perception method of network media devices is adopted to obtain the characteristic sequence of content to be played through the threat monitoring server, and compare it with the characteristic sequence of content currently played to determine whether the content is legal. The method includes the comparison of face and scene feature sequences, high-frequency central word collection sequences and audio fingerprint sequences. If it is found that they are illegal, it is determined that the device is in a threatened state.
It effectively improves the accuracy of real-time threat perception of online media devices, can promptly detect and respond to illegal playback threats, ensure content security, and reduce the risk of misleading and panic.
Smart Images

Figure CN120017302A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security technology, and in particular to a real-time threat perception method, device and electronic equipment for network media equipment. Background Art
[0002] In modern society, media devices such as electronic display screens have become indispensable information display tools and are widely used in commercial advertising, public information announcements, mass propaganda and education, and other fields. These media devices not only improve the efficiency and coverage of information dissemination, but also greatly enrich people's daily lives. With the development of technology, the application scenarios of network media devices are increasing. Most of them store playback content through cloud center resource servers, and network media playback devices in different locations obtain playback resources through the network for playback. This centralized and unified management method not only improves the update efficiency of playback content, but also reduces operating costs.
[0003] However, in public places, network media devices have become the main target of attacks by criminals due to their large number of display screens and wide audiences. Through network attacks, criminals can replace or tamper with the playback content of media devices, and then spread false information or conduct malicious propaganda, which will not only mislead the public, but also may cause social panic and instability. Therefore, the security issues of network media devices are becoming increasingly prominent, and how to ensure the normal operation and content security of these devices has become a technical problem that needs to be solved urgently.
[0004] Therefore, a method, device and electronic device for real-time threat perception of network media equipment are proposed. Summary of the invention
[0005] This specification provides a method, device and electronic device for real-time threat perception of network media devices, which can effectively improve the accuracy of real-time threat perception of network media devices.
[0006] This specification provides a real-time threat perception method for network media devices, including:
[0007] The threat monitoring server obtains the content to be played by the network media device through the resource server;
[0008] The threat monitoring server analyzes the content to be played to obtain a facial and scene feature sequence, a high-frequency central word set sequence, and an audio fingerprint sequence;
[0009] The threat perception device obtains the current playing content of the network media device, and parses the current playing content to obtain a facial and scene feature sequence of the current playing content, a high-frequency central word set sequence of the current playing content, and an audio fingerprint sequence of the current playing content;
[0010] The threat perception device compares the face and scene feature sequences of the currently played content with the face and scene feature sequences, the high-frequency core word set sequences of the currently played content with the high-frequency core word set sequences, and the audio fingerprint sequences of the currently played content with the audio fingerprint sequences, respectively, to determine whether the currently played content is legal;
[0011] When the threat perception device determines that the currently played content is illegal, it perceives that the network media device is in a threatened state.
[0012] Optionally, the threat monitoring server analyzes the content to be played to obtain a facial and scene feature sequence, including:
[0013] Performing frame processing on the content to be played to obtain a plurality of single-frame images and their sequence;
[0014] Performing face recognition on the single-frame images respectively to obtain a face feature set;
[0015] Performing scene recognition on the single-frame images respectively to obtain a scene feature set;
[0016] Comparing the number and features of faces in the face feature sets of two consecutive single-frame images one by one;
[0017] When the number of faces in the face feature sets of two consecutive single-frame images is equal and the face features are the same, it is determined that the faces of the two consecutive single-frame images are repeated;
[0018] Comparing the number of objects and scene features in the scene feature sets of two consecutive single-frame images one by one;
[0019] When the number of objects in the scene feature sets of two consecutive single-frame images is equal and the scene features are the same, it is determined that the scenes of the two consecutive single-frame images are repeated;
[0020] When two consecutive single-frame images have repeated faces and repeated scenes, deleting the single-frame image that is later in sequence;
[0021] The number of faces and facial features in the facial feature set of two consecutive single-frame images that are compared one by one are returned until all single-frame images are compared, thereby obtaining a facial and scene feature sequence.
[0022] Optionally, the threat monitoring server analyzes the content to be played to obtain a high-frequency central word set sequence, including:
[0023] Performing frame processing on the content to be played to obtain a plurality of single-frame images and their sequence;
[0024] Extracting text from the single-frame images respectively to obtain a set of high-frequency central words;
[0025] Comparing the high-frequency core word sets of two consecutive single-frame images one by one to obtain the high-frequency core word similarity of the two consecutive single-frame images;
[0026] Determining whether the similarity of the high-frequency central words of two consecutive single-frame images meets a preset similarity;
[0027] When the high-frequency central word similarity of two consecutive single-frame images meets a preset similarity, deleting the single-frame image that is later in the order of the single-frame images;
[0028] The high-frequency central word set of the two consecutive single-frame images that are compared one by one is returned to obtain the high-frequency central word similarity of the two consecutive single-frame images, until all single-frame images are compared to obtain a high-frequency central word set sequence.
[0029] Optionally, the threat monitoring server analyzes the content to be played to obtain an audio fingerprint sequence, including:
[0030] The to-be-played content is parsed using the Chromaprint open source audio fingerprint algorithm library to obtain an audio fingerprint sequence.
[0031] Optionally, the threat perception device compares the face and scene feature sequences of the currently played content with the face and scene feature sequences to determine whether the currently played content is legal, including:
[0032] Comparing the number of faces and facial features in the face and scene feature sequence of the currently played content with the number of faces and facial features in the face and scene feature sequence one by one;
[0033] When the number of faces in the face feature set in the face and scene feature sequence of the currently playing content is equal to the number of faces in the face and scene feature sequence, and the face features are the same, determining that the face feature sequence of the currently playing content is legal;
[0034] Comparing the number of objects and scene features in the scene feature sets of two consecutive single-frame images one by one;
[0035] When the number of objects in the face and scene feature sequence of the currently played content is equal to the number of objects in the face and scene feature sequence, and the scene features are the same, determining that the scene feature sequence of the currently played content is legal;
[0036] Whether the image content of the currently played content is legal is determined based on the facial feature sequence and the scene feature sequence.
[0037] Optionally, the threat perception device compares the high-frequency core word set sequence of the currently played content with the high-frequency core word set sequence, and the audio fingerprint sequence of the currently played content with the audio fingerprint sequence, respectively, to determine whether the currently played content is legal, including:
[0038] Determine whether the high-frequency core word of the high-frequency core word set sequence of the currently played content and the high-frequency core word of the high-frequency core word set sequence have a similarity that meets a preset similarity;
[0039] and,
[0040] Determine whether the audio fingerprint sequence of the currently played content overlaps with the audio fingerprint sequence.
[0041] Optionally, when the threat perception device determines that the currently played content is illegal, the network media device is perceived to be in a threatened state, including:
[0042] The threat perception device determines that the currently played content is illegal because any one of the image content is illegal, the high-frequency central word is illegal, and the audio is illegal;
[0043] Recording the number of times the currently played content is illegal, and determining whether the number of times the currently played content is illegal exceeds a preset number;
[0044] When the number of illegalities of the currently played content exceeds a preset number, it is determined that the network media device is in a threatened state.
[0045] This specification provides a real-time threat perception device for network media devices, including:
[0046] The threat monitoring server obtains the content to be played by the network media device through the resource server;
[0047] The threat monitoring server analyzes the content to be played to obtain a facial and scene feature sequence, a high-frequency central word set sequence, and an audio fingerprint sequence;
[0048] The threat perception device obtains the current playing content of the network media device, and parses the current playing content to obtain a facial and scene feature sequence of the current playing content, a high-frequency central word set sequence of the current playing content, and an audio fingerprint sequence of the current playing content;
[0049] The threat perception device compares the face and scene feature sequences of the currently played content with the face and scene feature sequences, the high-frequency core word set sequences of the currently played content with the high-frequency core word set sequences, and the audio fingerprint sequences of the currently played content with the audio fingerprint sequences, respectively, to determine whether the currently played content is legal;
[0050] When the threat perception device determines that the currently played content is illegal, it perceives that the network media device is in a threatened state.
[0051] Optionally, the threat monitoring server analyzes the content to be played to obtain a facial and scene feature sequence, including:
[0052] Performing frame processing on the content to be played to obtain a plurality of single-frame images and their sequence;
[0053] Performing face recognition on the single-frame images respectively to obtain a face feature set;
[0054] Performing scene recognition on the single-frame images respectively to obtain a scene feature set;
[0055] Comparing the number and features of faces in the face feature sets of two consecutive single-frame images one by one;
[0056] When the number of faces in the face feature sets of two consecutive single-frame images is equal and the face features are the same, it is determined that the faces of the two consecutive single-frame images are repeated;
[0057] Comparing the number of objects and scene features in the scene feature sets of two consecutive single-frame images one by one;
[0058] When the number of objects in the scene feature sets of two consecutive single-frame images is equal and the scene features are the same, it is determined that the scenes of the two consecutive single-frame images are repeated;
[0059] When two consecutive single-frame images have repeated faces and repeated scenes, deleting the single-frame image that is later in sequence;
[0060] The number of faces and facial features in the facial feature set of two consecutive single-frame images that are compared one by one are returned until all single-frame images are compared, thereby obtaining a facial and scene feature sequence.
[0061] Optionally, the threat monitoring server analyzes the content to be played to obtain a high-frequency central word set sequence, including:
[0062] Performing frame processing on the content to be played to obtain a plurality of single-frame images and their sequence;
[0063] Extracting text from the single-frame images respectively to obtain a set of high-frequency central words;
[0064] Comparing the high-frequency core word sets of two consecutive single-frame images one by one to obtain the high-frequency core word similarity of the two consecutive single-frame images;
[0065] Determining whether the similarity of the high-frequency central words of two consecutive single-frame images meets a preset similarity;
[0066] When the high-frequency central word similarity of two consecutive single-frame images meets a preset similarity, deleting the single-frame image that is later in the order of the single-frame images;
[0067] The high-frequency central word set of the two consecutive single-frame images that are compared one by one is returned to obtain the high-frequency central word similarity of the two consecutive single-frame images, until all single-frame images are compared to obtain a high-frequency central word set sequence.
[0068] Optionally, the threat monitoring server analyzes the content to be played to obtain an audio fingerprint sequence, including:
[0069] The to-be-played content is parsed using the Chromaprint open source audio fingerprint algorithm library to obtain an audio fingerprint sequence.
[0070] Optionally, the threat perception device compares the face and scene feature sequences of the currently played content with the face and scene feature sequences to determine whether the currently played content is legal, including:
[0071] Comparing the number of faces and facial features in the face and scene feature sequence of the currently played content with the number of faces and facial features in the face and scene feature sequence one by one;
[0072] When the number of faces in the face feature set in the face and scene feature sequence of the currently playing content is equal to the number of faces in the face and scene feature sequence, and the face features are the same, determining that the face feature sequence of the currently playing content is legal;
[0073] Comparing the number of objects and scene features in the scene feature sets of two consecutive single-frame images one by one;
[0074] When the number of objects in the face and scene feature sequence of the currently played content is equal to the number of objects in the face and scene feature sequence, and the scene features are the same, determining that the scene feature sequence of the currently played content is legal;
[0075] Whether the image content of the currently played content is legal is determined based on the facial feature sequence and the scene feature sequence.
[0076] Optionally, the threat perception device compares the high-frequency core word set sequence of the currently played content with the high-frequency core word set sequence, and the audio fingerprint sequence of the currently played content with the audio fingerprint sequence, respectively, to determine whether the currently played content is legal, including:
[0077] Determine whether the high-frequency core word of the high-frequency core word set sequence of the currently played content and the high-frequency core word of the high-frequency core word set sequence have a similarity that meets a preset similarity;
[0078] and,
[0079] Determine whether the audio fingerprint sequence of the currently played content overlaps with the audio fingerprint sequence.
[0080] Optionally, when the threat perception device determines that the currently played content is illegal, the network media device is perceived to be in a threatened state, including:
[0081] The threat perception device determines that the currently played content is illegal because any one of the image content is illegal, the high-frequency central word is illegal, and the audio is illegal;
[0082] Recording the number of times the currently played content is illegal, and determining whether the number of times the currently played content is illegal exceeds a preset number;
[0083] When the number of illegalities of the currently played content exceeds a preset number, it is determined that the network media device is in a threatened state.
[0084] This specification also provides an electronic device, wherein the electronic device includes:
[0085] processor; and,
[0086] A memory storing computer executable instructions, which when executed cause the processor to perform any of the above methods.
[0087] The present specification also provides a computer-readable storage medium, wherein the computer-readable storage medium stores one or more programs, and when the one or more programs are executed by a processor, any of the above methods is implemented.
[0088] In the present invention, by extracting data features of the content to be played in different time periods and using them as a detection benchmark for legitimacy, the edge device can fully perceive illegal playback threats such as replacement and tampering in the three dimensions of images, text, and audio. This method of directly judging illegal playback threats from the playback results does not require in-depth study of specific technical means, which greatly enhances the intuitiveness and accuracy of threat perception. In addition, the method adopts a structure that combines a cloud-based threat monitoring server with an edge device, and the time-consuming and high-computing-power data feature extraction work is handed over to the cloud for processing, while the edge device focuses on using its own computing resources for data analysis, effectively reducing the burden on the server. The edge device also has the ability to initiate threat perception detection tasks at random time intervals and can dynamically adjust the detection interval, which not only reduces the overall computational complexity, but also significantly improves the response speed of the network security system, ensuring timely discovery and real-time response to threats. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0090] Figure 1 A schematic diagram of the principle of a real-time threat perception method for a network media device provided in an embodiment of this specification;
[0091] Figure 2 A schematic diagram of the structure of a real-time threat perception device for a network media device provided in an embodiment of this specification;
[0092] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of this specification;
[0093] Figure 4 A schematic diagram of a computer-readable medium provided for an embodiment of this specification. DETAILED DESCRIPTION
[0094] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are only examples, and those skilled in the art can think of other obvious variations. The basic principles of the present invention defined in the following description can be applied to other embodiments, variations, improvements, equivalents, and other technical solutions that do not deviate from the spirit and scope of the present invention.
[0095] The following is combined with Figure 1-4 The exemplary embodiments of the present invention are described more fully. However, the exemplary embodiments can be implemented in various forms, and it should not be understood that the present invention is limited to the embodiments set forth herein. On the contrary, providing these exemplary embodiments can make the present invention more comprehensive and complete, and it is more convenient to fully convey the inventive concept to those skilled in the art. The same reference numerals in the figures represent the same or similar elements, components or parts, and thus their repeated description will be omitted.
[0096] Under the premise of being consistent with the technical concept of the present invention, the features, structures, characteristics or other details described in a specific embodiment do not exclude that they can be combined in one or more other embodiments in a suitable manner.
[0097] In the description of specific embodiments, the features, structures, characteristics or other details described in the present invention are intended to enable those skilled in the art to fully understand the embodiments. However, it does not exclude that those skilled in the art can practice the technical solutions of the present invention without one or more of the specific features, structures, characteristics or other details.
[0098] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0099] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0100] The term "and / or" or "and / or" includes all combinations of any one or more of the associated listed items.
[0101] Figure 1 A schematic diagram of a method for real-time threat perception of a network media device provided in an embodiment of this specification, the method may include:
[0102] S110: The threat monitoring server obtains the content to be played by the network media device through the resource server;
[0103] In the specific implementation of this specification, the main function of the resource server is to provide various network media devices with the resources required for playback, including video, images, and audio, etc., through network connections. At the same time, it is also responsible for establishing a playback content mapping table for network media devices, which records in detail the unique identification code of each device (such as MAC address), the category information of the display screen (such as usage and location classification), time period information, and specific playback content and links. With the help of this mapping table, the resource server can configure content for different categories of display screens in different time periods, thereby accurately controlling the playback and display content of network media devices, and realizing centralized and unified management and efficient scheduling of multiple network media devices.
[0104] The network media device sends a request to the resource server through the network connection. The request contains the unique identification code information of the device itself. After receiving the request, the resource server will retrieve the content to be played from the storage according to the current time period and the device's identification code. Subsequently, the network media device receives and plays the specified content to ensure that the correct information is displayed in the correct time period. This process realizes the precise matching and automatic management of content playback.
[0105] The main responsibility of the threat monitoring server is to ensure time synchronization with the resource server in the network, and to obtain the playback content mapping table of the network media device and all the content to be played from the resource server through the network connection. For each content to be played, the threat monitoring server will further extract its facial and scene feature sequence, high-frequency center word set sequence and audio fingerprint sequence. Subsequently, this information will be added to the playback content mapping table of the network media device to achieve more refined content management. In addition, the threat monitoring server will also subscribe to the playback content mapping table information in the resource server, so that when the mapping table content is added, modified or deleted, it can obtain these changes in real time, and update the facial and scene feature sequence, high-frequency center word set sequence and audio fingerprint sequence of the changed content according to the same process to ensure the accuracy and timeliness of the information.
[0106] S120: The threat monitoring server analyzes the content to be played to obtain a facial and scene feature sequence, a high-frequency core word set sequence, and an audio fingerprint sequence;
[0107] Optionally, the S120 includes:
[0108] Performing frame processing on the content to be played to obtain a plurality of single-frame images and their sequence;
[0109] Performing face recognition on the single-frame images respectively to obtain a face feature set;
[0110] Performing scene recognition on the single-frame images respectively to obtain a scene feature set;
[0111] Comparing the number and features of faces in the face feature sets of two consecutive single-frame images one by one;
[0112] When the number of faces in the face feature sets of two consecutive single-frame images is equal and the face features are the same, it is determined that the faces of the two consecutive single-frame images are repeated;
[0113] Comparing the number of objects and scene features in the scene feature sets of two consecutive single-frame images one by one;
[0114] When the number of objects in the scene feature sets of two consecutive single-frame images is equal and the scene features are the same, it is determined that the scenes of the two consecutive single-frame images are repeated;
[0115] When two consecutive single-frame images have repeated faces and repeated scenes, deleting the single-frame image that is later in sequence;
[0116] The number of faces and facial features in the facial feature set of two consecutive single-frame images that are compared one by one are returned until all single-frame images are compared, thereby obtaining a facial and scene feature sequence.
[0117] In a specific embodiment of the present specification, when processing the video or image content to be played, for any image frame in the video, face recognition technology is used to extract the face feature data of all faces in the image frame, and these data are combined into a face feature set. At the same time, the object and scene recognition technology in image recognition is used to obtain the object and scene feature data in the image frame to form an object and scene feature set. These two sets together constitute the face and scene feature sequence data. In order to optimize the data, deduplication processing is performed according to the order of the image frames: for any two consecutive frames of images, if the number of faces in their face feature sets is equal and the face features can be matched one-to-one, and the number of objects in the object and scene feature sets is also equal and the scene data can be matched one-to-one, then it is determined that the content of the two frames of images is repeated, and the feature data of the latter frame of the image is removed. Starting from the first frame and the second frame of the video, until the last frame, the above deduplication operation is repeatedly performed to finally obtain the face and scene feature sequence after the video content is refined.
[0118] Optionally, the S120 includes:
[0119] Performing frame processing on the content to be played to obtain a plurality of single-frame images and their sequence;
[0120] Extracting text from the single-frame images respectively to obtain a set of high-frequency central words;
[0121] Comparing the high-frequency core word sets of two consecutive single-frame images one by one to obtain the high-frequency core word similarity of the two consecutive single-frame images;
[0122] Determining whether the similarity of the high-frequency central words of two consecutive single-frame images meets a preset similarity;
[0123] When the high-frequency central word similarity of two consecutive single-frame images meets a preset similarity, deleting the single-frame image that is later in the order of the single-frame images;
[0124] The high-frequency central word set of the two consecutive single-frame images that are compared one by one is returned to obtain the high-frequency central word similarity of the two consecutive single-frame images, until all single-frame images are compared to obtain a high-frequency central word set sequence.
[0125] In a specific embodiment of the present specification, when processing the video or image content to be played, the image frames are first extracted from the video, and the general text recognition technology is applied to identify and extract the text and subtitle information in these image frames. Subsequently, the extracted text and subtitle information are subjected to text processing to extract the high-frequency core words in the text, and these high-frequency core words are formed into a set. According to the order of the image frames, the high-frequency core word sets of a plurality of consecutive image frames constitute a high-frequency core word set sequence. In order to optimize the data, the high-frequency core word set is further de-duplicated: for the high-frequency core word sets extracted from any two consecutive frames of images, if the set similarity between them is higher than a preset threshold, the two sets are determined to be repeated, and the word set corresponding to the text in the latter frame of the image is removed. The calculation of the set similarity is a key step to help accurately determine the repeatability of the two sets. Starting from the first frame and the second frame of the video, until the last frame, the above-mentioned de-duplication operation is repeatedly performed to finally obtain a high-frequency core word set sequence after the video content is optimized.
[0126] Optionally, the S120 includes:
[0127] The to-be-played content is parsed using the Chromaprint open source audio fingerprint algorithm library to obtain an audio fingerprint sequence.
[0128] In the specific implementation of this specification, in order to process the audio part of the content to be played, the audio file corresponding to the content is first obtained. Then, the audio file is processed using the open source audio fingerprint algorithm library Chromaprint to generate its unique audio fingerprint sequence. This step is intended to provide a unique and identifiable identifier for the audio content through audio fingerprint technology, so as to facilitate subsequent audio content identification, retrieval or copyright protection applications.
[0129] S130: The threat perception device obtains the current playing content of the network media device, and parses the current playing content to obtain a face and scene feature sequence of the current playing content, a high-frequency core word set sequence of the current playing content, and an audio fingerprint sequence of the current playing content;
[0130] In a specific implementation of the present specification, the threat perception device is connected to the network media device and the threat monitoring server. First, the threat perception device obtains the unique identification code information of the current display screen from the network media device, and then uses this information to initiate a request to the threat monitoring server to obtain the characteristic data of the content that the network media device corresponding to the identification code should play in the current period. These characteristic data cover the first feature sequence of the image, the high-frequency central word set sequence, and the audio fingerprint sequence. In addition, the threat perception device will also start the attack threat perception detection task of the network media device at a random time interval (the random value is greater than 0 and less than the preset value), and monitor the currently playing content through edge computing technology to determine whether there is illegal playback.
[0131] S140: The threat perception device compares the face and scene feature sequence of the currently played content with the face and scene feature sequence, the high-frequency core word set sequence of the currently played content with the high-frequency core word set sequence, and the audio fingerprint sequence of the currently played content with the audio fingerprint sequence, to determine whether the currently played content is legal;
[0132] In a specific implementation of the present specification, the threat perception device is connected to the network media device and can obtain the currently playing video frame in real time. Using face recognition technology, the threat perception device extracts the facial feature data of all faces from the video frame to form a facial feature set; at the same time, through the object and scene recognition function in the image recognition technology, the device also extracts the object and scene feature data in the image frame to form an object and scene feature set. These two types of feature sets together constitute the facial and scene feature description of the currently playing content. In addition, the device also uses general text recognition technology to identify and extract text and subtitle information in the image frame, and further processes these text information to extract high-frequency central words to form a high-frequency central word set of the currently playing content. At the same time, the device obtains the audio clip within the preset time since the detection task was started, and uses the Chromaprint open source audio fingerprint algorithm library to calculate the audio fingerprint of the audio clip. Finally, the edge computing device compares these feature data calculated by itself with the feature data provided by the threat monitoring server to determine whether the content currently played by the network media device is legal.
[0133] Optionally, the S140 includes:
[0134] Comparing the number of faces and facial features in the face and scene feature sequence of the currently played content with the number of faces and facial features in the face and scene feature sequence one by one;
[0135] When the number of faces in the face feature set in the face and scene feature sequence of the currently playing content is equal to the number of faces in the face and scene feature sequence, and the face features are the same, determining that the face feature sequence of the currently playing content is legal;
[0136] Comparing the number of objects and scene features in the scene feature sets of two consecutive single-frame images one by one;
[0137] When the number of objects in the face and scene feature sequence of the currently played content is equal to the number of objects in the face and scene feature sequence, and the scene features are the same, determining that the scene feature sequence of the currently played content is legal;
[0138] Whether the image content of the currently played content is legal is determined based on the facial feature sequence and the scene feature sequence.
[0139] In the specific implementation of this specification, if there is any sequence element, so that the face feature set in the element is equal to the face feature set in the feature data in terms of the number of faces, and each face feature can be matched one by one; at the same time, the object and scene feature set in the element is also equal to the object and scene feature set in the feature data in terms of the number of objects, and each object and scene feature can be matched one by one, then the image content is determined to be legal. On the contrary, if these two conditions cannot be met at the same time, the image content is determined to be illegal.
[0140] Optionally, the S140 includes:
[0141] Determine whether the high-frequency core word of the high-frequency core word set sequence of the currently played content and the high-frequency core word of the high-frequency core word set sequence have a similarity that meets a preset similarity;
[0142] and,
[0143] Determine whether the audio fingerprint sequence of the currently played content overlaps with the audio fingerprint sequence.
[0144] In a specific implementation of the present specification, if there is any sequence element whose high-frequency center word set has a similarity with a given high-frequency center word set that is higher than a preset threshold, then the high-frequency center word is determined to be legal. If this condition cannot be met, the high-frequency center word is determined to be illegal. The fingerprint sequence of the audio segment calculated previously is searched in the complete audio fingerprint sequence. If a matching fingerprint can be found in the sequence, the audio is immediately determined to be legal. On the contrary, if a matching fingerprint cannot be found, the audio is determined to be illegal.
[0145] S150: When the threat perception device determines that the currently played content is illegal, it perceives that the network media device is in a threatened state.
[0146] Optionally, the S150 includes:
[0147] The threat perception device determines that the currently played content is illegal because any one of the image content is illegal, the high-frequency central word is illegal, and the audio is illegal;
[0148] Recording the number of times the currently played content is illegal, and determining whether the number of times the currently played content is illegal exceeds a preset number;
[0149] When the number of illegalities of the currently played content exceeds a preset number, it is determined that the network media device is in a threatened state.
[0150] In a specific implementation of the present specification, once any one of the image content, high-frequency central words or audio is determined to be illegal, the playback status of the current network media device is immediately marked as a suspected threat, and the value of the random time interval is correspondingly reduced so that the next detection can be performed more frequently. If the cumulative number of suspected threats reaches or exceeds the preset value, the playback status of the current network media device will be determined to be illegal playback, that is, it is in a threatened state. The current playback content is turned off through the media device control unit (connected to the threat perception device), and the threat record information containing monitoring data such as the legitimacy of the image content, the legitimacy of the high-frequency central words, and the legitimacy of the audio is sent to the threat monitoring server. It is worth noting that the data interaction between the threat perception device and the threat monitoring server uses encrypted transmission to ensure data security and privacy protection.
[0151] In the present invention, by extracting data features of the content to be played in different time periods and using them as a detection benchmark for legitimacy, the edge device can fully perceive illegal playback threats such as replacement and tampering in the three dimensions of images, text, and audio. This method of directly judging illegal playback threats from the playback results does not require in-depth study of specific technical means, which greatly enhances the intuitiveness and accuracy of threat perception. In addition, the method adopts a structure that combines a cloud-based threat monitoring server with an edge device, and the time-consuming and high-computing-power data feature extraction work is handed over to the cloud for processing, while the edge device focuses on using its own computing resources for data analysis, effectively reducing the burden on the server. The edge device also has the ability to initiate threat perception detection tasks at random time intervals and can dynamically adjust the detection interval, which not only reduces the overall computational complexity, but also significantly improves the response speed of the network security system, ensuring timely discovery and real-time response to threats.
[0152] Figure 2 A schematic diagram of a real-time threat perception device for a network media device provided in an embodiment of this specification, the device may include:
[0153] The threat monitoring server obtains the content to be played by the network media device through the resource server;
[0154] The threat monitoring server analyzes the content to be played to obtain a facial and scene feature sequence, a high-frequency central word set sequence, and an audio fingerprint sequence;
[0155] The threat perception device obtains the current playing content of the network media device, and parses the current playing content to obtain a facial and scene feature sequence of the current playing content, a high-frequency central word set sequence of the current playing content, and an audio fingerprint sequence of the current playing content;
[0156] The threat perception device compares the face and scene feature sequences of the currently played content with the face and scene feature sequences, the high-frequency core word set sequences of the currently played content with the high-frequency core word set sequences, and the audio fingerprint sequences of the currently played content with the audio fingerprint sequences, respectively, to determine whether the currently played content is legal;
[0157] When the threat perception device determines that the currently played content is illegal, it perceives that the network media device is in a threatened state.
[0158] Optionally, the threat monitoring server analyzes the content to be played to obtain a facial and scene feature sequence, including:
[0159] Performing frame processing on the content to be played to obtain a plurality of single-frame images and their sequence;
[0160] Performing face recognition on the single-frame images respectively to obtain a face feature set;
[0161] Performing scene recognition on the single-frame images respectively to obtain a scene feature set;
[0162] Comparing the number and features of faces in the face feature sets of two consecutive single-frame images one by one;
[0163] When the number of faces in the face feature sets of two consecutive single-frame images is equal and the face features are the same, it is determined that the faces of the two consecutive single-frame images are repeated;
[0164] Comparing the number of objects and scene features in the scene feature sets of two consecutive single-frame images one by one;
[0165] When the number of objects in the scene feature sets of two consecutive single-frame images is equal and the scene features are the same, it is determined that the scenes of the two consecutive single-frame images are repeated;
[0166] When two consecutive single-frame images have repeated faces and repeated scenes, deleting the single-frame image that is later in sequence;
[0167] The number of faces and facial features in the facial feature set of two consecutive single-frame images that are compared one by one are returned until all single-frame images are compared, thereby obtaining a facial and scene feature sequence.
[0168] Optionally, the threat monitoring server analyzes the content to be played to obtain a high-frequency central word set sequence, including:
[0169] Performing frame processing on the content to be played to obtain a plurality of single-frame images and their sequence;
[0170] Extracting text from the single-frame images respectively to obtain a set of high-frequency central words;
[0171] Comparing the high-frequency core word sets of two consecutive single-frame images one by one to obtain the high-frequency core word similarity of the two consecutive single-frame images;
[0172] Determining whether the similarity of the high-frequency central words of two consecutive single-frame images meets a preset similarity;
[0173] When the high-frequency central word similarity of two consecutive single-frame images meets a preset similarity, deleting the single-frame image that is later in the order of the single-frame images;
[0174] The high-frequency central word set of the two consecutive single-frame images that are compared one by one is returned to obtain the high-frequency central word similarity of the two consecutive single-frame images, until all single-frame images are compared to obtain a high-frequency central word set sequence.
[0175] Optionally, the threat monitoring server analyzes the content to be played to obtain an audio fingerprint sequence, including:
[0176] The to-be-played content is parsed using the Chromaprint open source audio fingerprint algorithm library to obtain an audio fingerprint sequence.
[0177] Optionally, the threat perception device compares the face and scene feature sequences of the currently played content with the face and scene feature sequences to determine whether the currently played content is legal, including:
[0178] Comparing the number of faces and facial features in the face and scene feature sequence of the currently played content with the number of faces and facial features in the face and scene feature sequence one by one;
[0179] When the number of faces in the face feature set in the face and scene feature sequence of the currently playing content is equal to the number of faces in the face and scene feature sequence, and the face features are the same, determining that the face feature sequence of the currently playing content is legal;
[0180] Comparing the number of objects and scene features in the scene feature sets of two consecutive single-frame images one by one;
[0181] When the number of objects in the face and scene feature sequence of the currently played content is equal to the number of objects in the face and scene feature sequence, and the scene features are the same, determining that the scene feature sequence of the currently played content is legal;
[0182] Whether the image content of the currently played content is legal is determined based on the facial feature sequence and the scene feature sequence.
[0183] Optionally, the threat perception device compares the high-frequency core word set sequence of the currently played content with the high-frequency core word set sequence, and the audio fingerprint sequence of the currently played content with the audio fingerprint sequence, respectively, to determine whether the currently played content is legal, including:
[0184] Determine whether the high-frequency core word of the high-frequency core word set sequence of the currently played content and the high-frequency core word of the high-frequency core word set sequence have a similarity that meets a preset similarity;
[0185] and,
[0186] Determine whether the audio fingerprint sequence of the currently played content overlaps with the audio fingerprint sequence.
[0187] Optionally, when the threat perception device determines that the currently played content is illegal, the network media device is perceived to be in a threatened state, including:
[0188] The threat perception device determines that the currently played content is illegal because any one of the image content is illegal, the high-frequency central word is illegal, and the audio is illegal;
[0189] Recording the number of times the currently played content is illegal, and determining whether the number of times the currently played content is illegal exceeds a preset number;
[0190] When the number of illegalities of the currently played content exceeds a preset number, it is determined that the network media device is in a threatened state.
[0191] The functions of the device in the embodiment of the present invention have been described in the above method embodiment, so for details not provided in the description of this embodiment, please refer to the relevant description in the above embodiment, and no further description will be given here.
[0192] Based on the same inventive concept, an embodiment of this specification also provides an electronic device.
[0193] The following describes an electronic device embodiment of the present invention, which can be regarded as a specific physical implementation of the method and device embodiments of the present invention. The details described in the electronic device embodiment of the present invention should be regarded as a supplement to the above method or device embodiments; details not disclosed in the electronic device embodiment of the present invention can be implemented with reference to the above method or device embodiments.
[0194] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. Figure 3 The electronic device 300 according to the embodiment of the present invention is described. Figure 3 The electronic device 300 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0195] like Figure 3 As shown, the electronic device 300 is in the form of a general computing device. The components of the electronic device 300 may include, but are not limited to: at least one processing unit 310, at least one storage unit 320, a bus 330 connecting different system components (including the storage unit 320 and the processing unit 310), a display unit 340, etc.
[0196] The storage unit stores program codes, which can be executed by the processing unit 310, so that the processing unit 310 performs the steps according to various exemplary embodiments of the present invention described in the above processing method section of this specification. For example, the processing unit 310 can perform the following steps: Figure 1 Steps shown.
[0197] The storage unit 320 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 3201 and / or a cache memory unit 3202 , and may further include a read-only memory unit (ROM) 3203 .
[0198] The storage unit 320 may also include a program / utility 3204 having a set (at least one) of program modules 3205, such program modules 3205 including but not limited to: an operating system, one or more application programs, other program modules and program data, each of which or some combination may include the implementation of a network environment.
[0199] Bus 330 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0200] The electronic device 300 may also communicate with one or more external devices 400 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable viewers to interact with the electronic device 300, and / or any device that enables the electronic device 300 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed through an input / output (I / O) interface 350. Furthermore, the electronic device 300 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 360. The network adapter 360 may communicate with other modules of the electronic device 300 through the bus 330. It should be understood that although Figure 3 Not shown, other hardware and / or software modules may be used in conjunction with the electronic device 300, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0201] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the exemplary embodiments described in the present invention can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation method of the present invention can be embodied in the form of a software product, which can be stored in a computer-readable storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.) or on a network, including a number of instructions to enable a computing device (which can be a personal computer, server, or network device, etc.) to execute the above method according to the present invention. When the computer program is executed by a data processing device, the computer-readable medium can implement the above method of the present invention, that is: Figure 1 The method shown.
[0202] Figure 4 A schematic diagram of a computer-readable medium provided for an embodiment of this specification.
[0203] accomplish Figure 1The computer program of the method shown can be stored on one or more computer readable media. The computer readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0204] The computer readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by an instruction execution system, an apparatus, or a device or used in combination with it. The program code contained on the readable storage medium may be transmitted with any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0205] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the viewer computing device, partially on the viewer device, as a stand-alone software package, partially on the viewer computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the viewer computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0206] In summary, the present invention can be implemented in hardware, or in a software module running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that general data processing devices such as microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the embodiments of the present invention. The present invention can also be implemented as a device or apparatus program (e.g., a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0207] The specific embodiments described above further describe the purpose, technical solutions and beneficial effects of the present invention in detail. It should be understood that the present invention is not inherently related to any specific computer, virtual device or electronic device, and various general devices can also implement the present invention. The above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
[0208] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.
[0209] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A real-time threat perception method for network media devices, characterized in that: include: The threat monitoring server obtains the content to be played by the network media device through the resource server; The threat monitoring server analyzes the content to be played to obtain a facial and scene feature sequence, a high-frequency central word set sequence, and an audio fingerprint sequence; The threat perception device obtains the current playing content of the network media device, and parses the current playing content to obtain a facial and scene feature sequence of the current playing content, a high-frequency central word set sequence of the current playing content, and an audio fingerprint sequence of the current playing content; The threat perception device compares the face and scene feature sequences of the currently played content with the face and scene feature sequences, the high-frequency core word set sequences of the currently played content with the high-frequency core word set sequences, and the audio fingerprint sequences of the currently played content with the audio fingerprint sequences, respectively, to determine whether the currently played content is legal; When the threat perception device determines that the currently played content is illegal, it perceives that the network media device is in a threatened state.
2. The real-time threat perception method for network media devices according to claim 1, characterized in that: The threat monitoring server analyzes the content to be played to obtain a facial and scene feature sequence, including: Performing frame processing on the content to be played to obtain a plurality of single-frame images and their sequence; Performing face recognition on the single-frame images respectively to obtain a face feature set; Performing scene recognition on the single-frame images respectively to obtain a scene feature set; Comparing the number and features of faces in the face feature sets of two consecutive single-frame images one by one; When the number of faces in the face feature sets of two consecutive single-frame images is equal and the face features are the same, it is determined that the faces of the two consecutive single-frame images are repeated; Comparing the number of objects and scene features in the scene feature sets of two consecutive single-frame images one by one; When the number of objects in the scene feature sets of two consecutive single-frame images is equal and the scene features are the same, it is determined that the scenes of the two consecutive single-frame images are repeated; When two consecutive single-frame images have repeated faces and repeated scenes, deleting the single-frame image that is later in sequence; The number of faces and facial features in the facial feature set of two consecutive single-frame images that are compared one by one are returned until all single-frame images are compared, thereby obtaining a facial and scene feature sequence.
3. The real-time threat perception method for network media devices according to claim 2, characterized in that: The threat monitoring server analyzes the content to be played to obtain a high-frequency core word set sequence, including: Performing frame processing on the content to be played to obtain a plurality of single-frame images and their sequence; Extracting text from the single-frame images respectively to obtain a set of high-frequency central words; Comparing the high-frequency core word sets of two consecutive single-frame images one by one to obtain the high-frequency core word similarity of the two consecutive single-frame images; Determining whether the similarity of the high-frequency central words of two consecutive single-frame images meets a preset similarity; When the high-frequency central word similarity of two consecutive single-frame images meets a preset similarity, deleting the single-frame image that is later in the order of the single-frame images; The high-frequency central word set of the two consecutive single-frame images that are compared one by one is returned to obtain the high-frequency central word similarity of the two consecutive single-frame images, until all single-frame images are compared to obtain a high-frequency central word set sequence.
4. The real-time threat perception method for network media devices according to claim 3, characterized in that: The threat monitoring server analyzes the content to be played to obtain an audio fingerprint sequence, including: The to-be-played content is parsed using the Chromaprint open source audio fingerprint algorithm library to obtain an audio fingerprint sequence.
5. The real-time threat perception method for network media devices according to claim 4, characterized in that: The threat perception device compares the face and scene feature sequences of the currently played content with the face and scene feature sequences to determine whether the currently played content is legal, including: Comparing the number of faces and facial features in the face and scene feature sequence of the currently played content with the number of faces and facial features in the face and scene feature sequence one by one; When the number of faces in the face feature set in the face and scene feature sequence of the currently playing content is equal to the number of faces in the face and scene feature sequence, and the face features are the same, determining that the face feature sequence of the currently playing content is legal; Comparing the number of objects and scene features in the scene feature sets of two consecutive single-frame images one by one; When the number of objects in the face and scene feature sequence of the currently played content is equal to the number of objects in the face and scene feature sequence, and the scene features are the same, determining that the scene feature sequence of the currently played content is legal; Whether the image content of the currently played content is legal is determined based on the facial feature sequence and the scene feature sequence.
6. The real-time threat perception method for network media devices according to claim 5, characterized in that: The threat perception device compares the high-frequency core word set sequence of the currently played content with the high-frequency core word set sequence, and the audio fingerprint sequence of the currently played content with the audio fingerprint sequence, respectively, to determine whether the currently played content is legal, including: Determine whether the high-frequency core word of the high-frequency core word set sequence of the currently played content and the high-frequency core word of the high-frequency core word set sequence have a similarity that meets a preset similarity; and, Determine whether the audio fingerprint sequence of the currently played content overlaps with the audio fingerprint sequence.
7. The real-time threat perception method for network media devices according to claim 6, characterized in that: When the threat perception device determines that the currently played content is illegal, the network media device is perceived to be in a threatened state, including: The threat perception device determines that the currently played content is illegal because any one of the image content is illegal, the high-frequency central word is illegal, and the audio is illegal; Recording the number of times the currently played content is illegal, and determining whether the number of times the currently played content is illegal exceeds a preset number; When the number of illegalities of the currently played content exceeds a preset number, it is determined that the network media device is in a threatened state.
8. A real-time threat perception device for network media equipment, characterized in that: include: The threat monitoring server obtains the content to be played by the network media device through the resource server; The threat monitoring server analyzes the content to be played to obtain a facial and scene feature sequence, a high-frequency central word set sequence, and an audio fingerprint sequence; The threat perception device obtains the current playing content of the network media device, and parses the current playing content to obtain a facial and scene feature sequence of the current playing content, a high-frequency central word set sequence of the current playing content, and an audio fingerprint sequence of the current playing content; The threat perception device compares the face and scene feature sequences of the currently played content with the face and scene feature sequences, the high-frequency core word set sequences of the currently played content with the high-frequency core word set sequences, and the audio fingerprint sequences of the currently played content with the audio fingerprint sequences, respectively, to determine whether the currently played content is legal; When the threat perception device determines that the currently played content is illegal, it perceives that the network media device is in a threatened state.
9. An electronic device, wherein: The electronic device includes: processor; and, A memory storing computer executable instructions which, when executed, cause the processor to perform a method according to any one of claims 1-7.
10. A computer-readable storage medium, wherein: The computer-readable storage medium stores one or more programs, and when the one or more programs are executed by a processor, the method of any one of claims 1 to 7 is implemented.