A violent and terrorist video recognition method based on key feature fusion
By constructing a key feature database and employing synchronous spatiotemporal fusion technology, the high false alarm rate and source tracing difficulties of existing violent and terrorist video recognition algorithms have been solved, enabling accurate identification and source tracing of violent and terrorist videos and preventing their spread.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING WEIZHIXINYE TECH CO LTD
- Filing Date
- 2023-09-20
- Publication Date
- 2026-05-19
AI Technical Summary
Existing violent and terrorist video recognition algorithms have a high false alarm rate when detecting violent and terrorist videos, especially when some segments contain few violent and terrorist elements but have a large impact. Furthermore, existing technologies cannot trace back to the original video.
By constructing a method for identifying violent and terrorist videos based on key feature fusion, including pre-building a key feature database, extracting audio and image features of the video segments to be identified, performing synchronous spatiotemporal fusion and tracing the propagation path, determining the video subject, and identifying whether the video is a violent and terrorist video.
It enables accurate identification and tracing of violent and terrorist videos, preventing their spread at the source and improving the accuracy and efficiency of identification.
Smart Images

Figure CN117315340B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of violent and terrorist video recognition technology, and in particular to a method for violent and terrorist video recognition based on key feature fusion. Background Technology
[0002] Currently, using artificial intelligence algorithms to automate the review of massive amounts of data on internet platforms has become the most widely used technology in the field of content moderation. Compared with traditional manual review methods, automated review not only greatly improves the accuracy of review for internet companies but also reduces operating costs. Terrorist content is a common type of harmful information explicitly defined by the state. Today, with the development of artificial intelligence, the original manual review stage in enterprises has been gradually replaced by various algorithms. With massive amounts of data, the advantages of automated review are very obvious. However, it is undeniable that even with increasingly higher accuracy, existing algorithms still face challenges in completely eliminating manual review and replacing it entirely with machines. Video, as one of the most popular forms of mass media, often contains content related to violence and terrorism. Due to the temporal relationship between frames in videos and the ambiguity in defining violent and terrorist content, algorithms still have a high false positive rate when detecting such videos. In recent years, research on algorithms for detecting violent and terrorist videos has increased, but most studies focus on analyzing video frames. Besides keyframes, the temporal relationship between audio and video frames is also crucial in the modeling process. Even during manual review, reviewers need to combine audio and video content to make accurate judgments.
[0003] With the emergence and development of computers and related technologies, algorithms have far surpassed manual methods in many scenarios for classifying and detecting data such as videos and images.
[0004] However, in the identification of violent and terrorist videos, it is currently only possible to identify complete violent and terrorist videos. Some violent and terrorist videos contain relatively few violent and terrorist elements, but after attracting users, users can trace the source or be recommended by many algorithms. The original video is a violent and terrorist video, but some of its segments contain very few violent and terrorist elements, making it difficult to identify. Summary of the Invention
[0005] This invention provides a method for identifying violent and terrorist videos based on key feature fusion, in order to address the aforementioned situation in the background technology.
[0006] This application proposes a method for identifying violent and terrorist videos based on key feature fusion, including:
[0007] Pre-build a database of key features based on the identification of violent and terrorist videos;
[0008] Acquire the video segment to be identified, and extract its content features and distribution channel data; among them,
[0009] Content features include audio features and image features;
[0010] Based on the key feature database, the audio key features and image key features of the video segment to be identified are spatiotemporally fused synchronously on the time axis to generate synchronous spatiotemporal features;
[0011] Based on the data from the dissemination channels, the dissemination path of the video to be identified is traced and the spatiotemporal features are simultaneously authenticated to determine the video subject of the video to be identified;
[0012] Based on the main body of the video, the video is identified as a violent and terrorist video, and when the main body of the video is a violent and terrorist video, the video segment to be identified is identified as a violent and terrorist video.
[0013] Preferably, the key feature database includes an audio key feature database and a video key feature database; wherein,
[0014] The audio key feature database includes: a scene audio key feature database, an event audio key feature database, and an object audio key feature database; among them,
[0015] The scene audio key feature database is constructed based on the event audio key feature database and the object audio key feature database;
[0016] The video key feature database includes a scene video key feature database, an event video key feature database, and an object video key feature database.
[0017] Preferably, the pre-construction of the key feature database based on violent and terrorist videos includes the following steps:
[0018] The design includes dynamic terrorist scenarios in two types: violent terrorist videos and violent terrorist audio.
[0019] The simulated terrorist behavior of participants in a dynamic terrorist scenario was conducted through audio text simulation and video image representation simulation.
[0020] Based on simulated terrorist behavior, key audio and video features of different participants were extracted.
[0021] The key features of audio and video are further subdivided into three types of terrorist risk factors: scene, event, and object. The key features of video and audio under the corresponding terrorist risk factors are determined, thereby constructing a key feature database containing terrorist elements.
[0022] Preferably, the content features of the video segment to be identified are extracted by a content capture mechanism; wherein,
[0023] The content capture mechanism is used to capture text, objects, people, and scene environment based on the frame images of the video segment to be identified;
[0024] The content capture mechanism is also used to mark terrorist elements based on perceived terrorist risks after extracting content features, and to record time-varying information of these elements; among other things,
[0025] Time-varying information and terrorist element markers are matched and aligned in time.
[0026] Preferably, the dissemination channel data includes the direct playback platform and dissemination path of the video segment to be identified; wherein,
[0027] Direct playback platforms include the first platform, the second platform, and the third platform; among them,
[0028] The first platform is a homogeneous content playback platform, used to determine the playback platform whose homogeneity with the video segment to be identified reaches a preset value;
[0029] The second platform is a content playback platform, used to play the video segments that are identified as the parts of the video to be identified;
[0030] The third platform is the secondary video playback platform, used to determine the playback platform that includes all content of the video segment to be identified;
[0031] The propagation path includes the main propagation path and propagation branch paths;
[0032] The main propagation path is a sequential path composed of the sequential propagation nodes of the video segment to be identified;
[0033] The propagation branch path is the propagation path associated with each node in the sequential propagation nodes.
[0034] Preferably, the synchronous spatiotemporal fusion includes:
[0035] Construct a timeline based on the video segment to be identified, and record frame images and audio data at each moment through the timeline;
[0036] The frame image and audio data are detected by the homomorphic feature detection model configured in the key feature database to determine whether homomorphic features exist.
[0037] When homomorphic features exist, they are used as key features of the target.
[0038] Based on the key features of the target, feature sorting is performed on a time axis to generate a feature sequence;
[0039] Based on the feature sequence, determine the sequence bits that simultaneously possess both key image features and key audio features at the same time, and form sequence pairs;
[0040] Based on the sequence pairs, spatiotemporal weighted fusion is performed to generate synchronous spatiotemporal features.
[0041] Preferably, the propagation path tracing includes:
[0042] Receive data on the distribution channels of the video segment to be identified;
[0043] By comparing the data from the dissemination channels with the preset sample data, a source tracing operation is performed to determine the video metadata;
[0044] Based on the video metadata, all publishing platforms of the video segment to be identified are determined, and the video subject of the earliest video segment to be identified among all publishing platforms is determined, thereby determining the platform from which the corresponding video was disseminated.
[0045] Preferably, the synchronous spatiotemporal feature authentication includes:
[0046] Based on data from distribution channels, obtain related videos of the video segments to be identified on different platforms;
[0047] The identifier axis is set according to the start and end image frames of the video segment to be identified, at a preset interval.
[0048] Each associated video is authenticated using a corresponding identifier axis;
[0049] Based on the corresponding authentication results, determine the set of target videos that may become the main subject of the associated videos.
[0050] Preferably, the method further includes:
[0051] Obtain multiple video frames from the video segment to be identified;
[0052] Determine the terrorism parameters corresponding to each of multiple video frames; among them,
[0053] Terrorism parameters are determined through key feature data;
[0054] The terrorism parameter is used to identify the weight of each video frame in the video segment to be identified;
[0055] Multiple video frames are fused based on their respective terrorism parameters to obtain a fused frame that associates the video segment to be identified with terrorism.
[0056] Preferably, the method further includes:
[0057] Obtain assessment data for violent and terrorist videos;
[0058] Based on the assessment data, hazard indicators and potential accident indicators are set for different key features on the video segments to be identified.
[0059] Based on probability weights, calculate the first weight of each indicator in the terrorism hazard index and event probability index;
[0060] Based on the scope weight, calculate the second weight of each indicator in the terrorist harm index and the event probability index;
[0061] The final weights are averaged based on the first and second weights.
[0062] The severity of the terrorist incident in the video segment to be identified is assessed based on the final weight.
[0063] The beneficial effects of this application are:
[0064] This invention does not directly judge the video itself when identifying violent and terrorist videos. Instead, it traces the source of the video to determine whether the original video is a violent and terrorist video, thereby deleting a series of related videos and preventing the spread of violent and terrorist videos from the source.
[0065] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.
[0066] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0067] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0068] Figure 1 This is a flowchart of a method for identifying violent and terrorist videos based on key feature fusion in an embodiment of the present invention;
[0069] Figure 2 This is a flowchart of synchronous spatiotemporal authentication in an embodiment of the present invention;
[0070] Figure 3 This is a flowchart illustrating the assessment of the degree of harm caused by videos in an embodiment of the present invention. Detailed Implementation
[0071] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0072] This application proposes a method for identifying violent and terrorist videos based on key feature fusion, including:
[0073] Pre-build a database of key features based on the identification of violent and terrorist videos;
[0074] Acquire the video segment to be identified, and extract its content features and distribution channel data; among them,
[0075] Content features include audio features and image features;
[0076] Based on the key feature database, the audio key features and image key features of the video segment to be identified are spatiotemporally fused synchronously on the time axis to generate synchronous spatiotemporal features;
[0077] Based on the data from the dissemination channels, the dissemination path of the video to be identified is traced and the spatiotemporal features are simultaneously authenticated to determine the video subject of the video to be identified;
[0078] Based on the main body of the video, the video is identified as a violent and terrorist video, and when the main body of the video is a violent and terrorist video, the video segment to be identified is identified as a violent and terrorist video.
[0079] The principle behind the above technical solution is as follows:
[0080] As attached Figure 1 As shown, the identification of violent terrorist videos in this invention is mainly based on the identification of inconspicuous violent terrorist videos. These violent terrorist videos often do not contain violent terrorist elements, but the original video is a violent terrorist video.
[0081] The key feature database is a database built based on terrorist elements in terrorist videos, including terrorist audio and text, terrorist objects such as shells and pistols, and terrorist images or images of terrorist actions.
[0082] The content features are the text content, image content, and audio content corresponding to the video to be identified;
[0083] The data on dissemination channels includes the platforms through which the video was disseminated, the platforms on which it was reposted, and the identity information of the individuals who reposted it.
[0084] The propagation path tracing and synchronous spatiotemporal feature authentication involves tracing the propagation path of the video to be identified to determine all related videos and the entities that published them, i.e., videos from the same batch.
[0085] Synchronous spatiotemporal authentication is designed to ensure that the video subject includes the video to be identified, and that it is the original and most complete video.
[0086] By analyzing the most comprehensive video footage, we can determine whether it is a violent or terrorist video.
[0087] The beneficial effects of the above technical solution are as follows:
[0088] This invention does not directly judge the video itself when identifying violent and terrorist videos. Instead, it traces the source of the video to determine whether the original video is a violent and terrorist video, thereby deleting a series of related videos and preventing the spread of violent and terrorist videos from the source.
[0089] Specifically, the key feature database includes an audio key feature database and a video key feature database; wherein,
[0090] The audio key feature database includes: a scene audio key feature database, an event audio key feature database, and an object audio key feature database; among them,
[0091] The scene audio key feature database is constructed based on the event audio key feature database and the object audio key feature database;
[0092] The video key feature database includes a scene video key feature database, an event video key feature database, and an object video key feature database.
[0093] The principle behind the above technical solution is as follows:
[0094] The key feature database includes audio key features and video key features. The two types of features include feature data of terrorism in different scenarios, feature data of different events, and feature data when different objects are present. The objects are objects that exist in the context of terrorism, mainly corresponding to weapons. The scenarios are the scene and location features in the video or audio that clearly indicate the possible occurrence of terrorism. This includes all features from both audio and video that can characterize the existence or possibility of terrorism.
[0095] The beneficial effects of the above technical solution are as follows:
[0096] The key feature database of this application can determine the key features of violent and terrorist videos based on audio and video from three aspects: scene, time, and object.
[0097] Specifically, the pre-construction of a key feature database based on violent and terrorist videos includes the following steps:
[0098] The design includes dynamic terrorist scenarios in two types: violent terrorist videos and violent terrorist audio.
[0099] The simulated terrorist behavior of participants in a dynamic terrorist scenario was conducted through audio text simulation and video image representation simulation.
[0100] Based on simulated terrorist behavior, key audio and video features of different participants were extracted.
[0101] The key features of audio and video are further subdivided into three types of terrorist risk factors: scene, event, and object. The key features of video and audio under the corresponding terrorist risk factors are determined, thereby constructing a key feature database containing terrorist elements.
[0102] The principle behind the above technical solution is as follows:
[0103] In the process of building the database, this invention simulates terrorist scenarios based on three factors: scene, event, and object. Through the simulation, it extracts key factors in broadcast control behavior and then constructs the database.
[0104] Audio text simulation can simulate audio information in violent terrorist videos, while video image representation simulation can simulate all possible scenarios where violent terrorist events may occur. Furthermore, through simulated violent terrorist behavior, violent terrorist videos can be subdivided, and violent terrorist features from different angles can be extracted according to the subdivision to generate a violent terrorist feature database.
[0105] The beneficial effects of the above technical solution are as follows:
[0106] The key feature database of this application can collect audio and video features of violent terrorist videos from three perspectives: scene, time, and object. It adopts a simulation method, thus highlighting more characteristics of violent terrorist phenomena.
[0107] Specifically, the content features of the video segment to be identified are extracted by a content capture mechanism; wherein,
[0108] The content capture mechanism is used to capture text, objects, people, and scene environment based on the frame images of the video segment to be identified;
[0109] The content capture mechanism is also used to mark terrorist elements based on perceived terrorist risks after extracting content features, and to record time-varying information of these elements; among other things,
[0110] Time-varying information and terrorist element markers are matched and aligned in time.
[0111] The principle behind the above technical solution is as follows:
[0112] In this embodiment, the content capture mechanism is used to capture content in the video to be identified that may contain violent and terrorist elements. After capturing violent and terrorist elements, the mechanism predicts the potential risks, i.e., the cognitive risks of violence and terrorism, based on the violent and terrorist elements, and then marks them accordingly. Within a video segment, the mechanism determines whether each marked violent and terrorist element has changed over time, thereby performing temporal matching to achieve content capture, marking, and change tracking of violent and terrorist elements.
[0113] The beneficial effects of the above technical solution are as follows:
[0114] This application enables the capture and tracking of violent and terrorist elements, allowing for rapid analysis of different violent and terrorist elements. It can also capture complete videos of the same type more quickly based on clips.
[0115] Specifically, the dissemination channel data includes the direct playback platforms and dissemination paths of the video segment to be identified; wherein,
[0116] Direct playback platforms include the first platform, the second platform, and the third platform; among them,
[0117] The first platform is a homogeneous content playback platform, used to determine the playback platform whose homogeneity with the video segment to be identified reaches a preset value;
[0118] The second platform is a content playback platform, used to play the video segments that are identified as the parts of the video to be identified;
[0119] The third platform is the secondary video playback platform, used to determine the playback platform that includes all content of the video segment to be identified;
[0120] The propagation path includes the main propagation path and propagation branch paths;
[0121] The main propagation path is a sequential path composed of the sequential propagation nodes of the video segment to be identified;
[0122] The propagation branch path is the propagation path associated with each node in the sequential propagation nodes.
[0123] The principle behind the above technical solution is as follows:
[0124] In the process of channel identification, this application, in order to trace the source of video clips, will look for the specific playback platforms of the video clips. Therefore, this application proposes three platforms: a homogenization platform, where, in the complete video of a homogenization platform, if there is a video with a very high similarity to the video clip to be identified, it belongs to the first platform. This similarity includes the entirety of the complete video or a part of the complete video being very similar to the video clip to be identified. If the video to be identified is a cropped video of a complete video, after tracing its source, it belongs to a part of another cropped video. Then, through multiple tracing, the original complete video is found. At this time, the multiple playback platforms of video clips that are only cropped videos encountered in the multiple tracing process constitute the second platform. The third platform is the main video playback platform, that is, the playback platform of the complete original video of the video clip to be identified. In this tracing process, the tracing path of the video to be disseminated is also determined, thereby determining the entire dissemination path of violent and terrorist videos in the process of identifying violent and terrorist videos, so as to achieve comprehensive deletion of violent and terrorist videos across the entire network when they are discovered.
[0125] Specifically, the synchronous spatiotemporal fusion includes:
[0126] Construct a timeline based on the video segment to be identified, and record frame images and audio data at each moment through the timeline;
[0127] The frame image and audio data are detected by the homomorphic feature detection model configured in the key feature database to determine whether homomorphic features exist.
[0128] When homomorphic features exist, they are used as key features of the target.
[0129] Based on the key features of the target, feature sorting is performed on a time axis to generate a feature sequence;
[0130] Based on the feature sequence, determine the sequence bits that simultaneously possess both key image features and key audio features at the same time, and form sequence pairs;
[0131] Based on the sequence pairs, spatiotemporal weighted fusion is performed to generate synchronous spatiotemporal features.
[0132] The principle behind the above technical solution is as follows:
[0133] As attached Figure 2 As shown, in the process of synchronous spatiotemporal fusion of violent terrorist elements in the video to be identified, this application synchronously fuses the audio data and video data at the moment when the violent terrorist element is generated, thereby ensuring the accuracy of identification and tracing during the fusion of the video segment to be identified and the initial video.
[0134] By constructing a timeline, audio and video data at each moment can be extracted. Then, using a homomorphic feature detection model in the key feature database, the presence of synchronized audio and video data containing terrorist elements can be detected. Furthermore, when homomorphic features are present, these features are used as the primary key features for tracing the source of the video segment to be identified—that is, the target key features. Through a series of target key features, corresponding feature sequences can be generated. These feature sequences are then used to generate audio and video sequence pairs. Finally, through spatiotemporal weighted fusion of the sequence pairs, the most easily identifiable synchronized spatiotemporal features in each sequence pair are determined.
[0135] The beneficial effects of the above technical solution are as follows:
[0136] This application uses a synchronous spatiotemporal fusion processing method to synchronously queue and sort audio and video data of different terrorist elements. This synchronous spatiotemporal fusion method makes it faster to trace the source and find the spread path of terrorist elements through video.
[0137] Specifically, the transmission path is traced back to terrorism:
[0138] Receive data on the distribution channels of the video segment to be identified;
[0139] By comparing preset data from the dissemination channel data with pre-processed sample data, a source tracing operation is performed to determine the video metadata.
[0140] Based on the video metadata, all publishing platforms of the video segment to be identified are determined, and the video subject of the earliest video segment to be identified among all publishing platforms is determined, thereby determining the platform from which the corresponding video was disseminated.
[0141] The principle behind the above technical solution is as follows:
[0142] In this embodiment, for the source tracing process of the propagation path, firstly, based on the video segment to be identified, the real-time propagation channel data of the video is obtained. Then, based on the real-time propagation channel data, sample data is a metadata representation of channel data. Thus, through this comparison, the metadata collection of propagation channel data can be achieved, which is a comparative source tracing operation.
[0143] Then, based on the video metadata, all platforms that published the video segment to be identified can be identified. By judging the complete video from the publishing platforms and determining the earliest video, the platform from which the spread originated can be determined.
[0144] The beneficial effects of the above technical solution are as follows:
[0145] This application enables full-network tracing of the source of a video segment to be identified by determining the platform from which it originated, thereby pinpointing the platform that first published the complete video containing the segment. In this process, the dissemination channel data includes metadata from the transmission process of the video segment. This metadata is then compared with pre-set sample data to determine the corresponding video metadata. Furthermore, the metadata is used to identify all platforms that published the video segment, thus achieving source tracing of the dissemination path.
[0146] The beneficial effects of the above technical solution are as follows:
[0147] This application can trace the source of data from the dissemination channels to determine the specific dissemination platform of the video segment to be identified, thereby achieving source tracing and positioning.
[0148] Specifically, the synchronous spatiotemporal feature authentication includes:
[0149] Based on data from distribution channels, obtain related videos of the video segments to be identified on different platforms;
[0150] The identifier axis is set according to the start and end image frames of the video segment to be identified, at a preset interval.
[0151] Each associated video is authenticated sequentially based on the identifier axis, using the video containing the segment to be identified.
[0152] Based on the corresponding authentication results, determine the set of target videos that may become the main subject of the associated videos.
[0153] The principle behind the above technical solution is as follows:
[0154] This invention, through synchronous spatiotemporal authentication, can determine all dissemination channels and identical videos within the same platform that share the same source, or the preceding and following segments of the video to be identified, generate a corresponding set, and then perform comprehensive deletion.
[0155] In this synchronous spatiotemporal feature authentication process, the propagation channel data can determine the location of related videos on different platforms and the video segment to be identified. After location, the feature distribution of the video segment to be identified is determined by the start image frame and the end image frame. Then, the video features corresponding to each moment are recorded on the label axis through the feature distribution, thereby realizing the corresponding identification and authentication. When the identification and authentication result shows that the related video contains the video segment to be identified, it means that the related video can be used as the target video.
[0156] The beneficial effects of the above technical solution are as follows:
[0157] This application can determine a set of target videos that may be video sources of the video segment to be identified by associating the video segments to be identified with the associated videos. In this process, this application adopts an interval time identifier axis, which can compare with associated videos based on the correspondence between time and video features. It has an authentication method that combines time authentication and feature authentication with symmetrical association. At the same time, it can save time, as it is not necessary to compare every feature, and it can also speed up the review efficiency of associated food products.
[0158] Specifically, the method further includes:
[0159] Obtain multiple video frames from the video segment to be identified;
[0160] Determine the terrorism parameters corresponding to each of multiple video frames; among them,
[0161] Terrorism parameters are determined through key feature data;
[0162] The terrorism parameter is used to identify the weight of each video frame in the video segment to be identified;
[0163] Multiple video frames are fused based on their respective terrorism parameters to obtain a fused frame that associates the video segment to be identified with terrorism.
[0164] The principle behind the above technical solution is as follows:
[0165] This application addresses the challenge of locating different terrorist videos when identifying segments as such. The process involves calculating terrorist parameters from multiple video frames, characterizing the terrorist elements present in each frame. Based on the weights of each frame, the corresponding terrorist elements are determined. A fusion frame then attempts to correlate and fuse different terrorist elements to identify the specific terrorist event. Therefore, the fusion frame can quickly determine the corresponding terrorist behavior when terrorist elements are present.
[0166] The beneficial effects of the above technical solution are as follows:
[0167] This application can quickly determine the corresponding terrorist behavior in the video clip, thereby identifying terrorist events. In this process, because the weight of video frames is used to perform video frame fusion, a terrorist event can be quickly spliced into a specific behavior judgment event.
[0168] Specifically, the method further includes:
[0169] Obtain assessment data for violent and terrorist videos;
[0170] Based on the assessment data, hazard indicators and potential accident indicators are set for different key features on the video segments to be identified.
[0171] Based on probability weights, calculate the first weight of each indicator in the terrorism hazard index and event probability index;
[0172] Based on the scope weight, calculate the second weight of each indicator in the terrorist harm index and the event probability index;
[0173] The final weights are averaged based on the first and second weights.
[0174] The severity of the terrorist incident in the video segment to be identified is assessed based on the final weight. The principle behind the above technical solution is as follows:
[0175] As attached Figure 3 As shown, this invention identifies dangers and accidents by combining different key features on the video segment to be identified with the original video.
[0176] In this determination process:
[0177] The first step is to evaluate the data, which includes data on different terrorist behaviors and terrorist characteristics in the video segments to be identified. Based on this data, corresponding evaluation indicators are set to assess the probability and specific harm of these terrorist behaviors and characteristics.
[0178] In its evaluation process, this application employs a weighted fusion and equalization method. The first weight is the probability of all occurrences, assessing the probability of terrorist acts and characteristics transforming into specific events. The second weight is a range weight, used to assess the specific range of potential impact of terrorist acts and events. This equalization process determines an average value, and based on this average value, the degree of harm of the terrorist event in the video segment to be identified is determined.
[0179] The beneficial effects of the above technical solution are as follows:
[0180] This application uses the weighting process from two different perspectives to determine the impact of terrorist incidents in terms of probability and scope, thereby determining the specific degree of harm.
[0181] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for identifying violent and terrorist videos based on key feature fusion, characterized in that, include: Pre-build a database of key features based on the identification of violent and terrorist videos; Acquire the video segment to be identified, and extract its content features and dissemination channel data. in, Content features include audio features and image features; Based on the key feature database, the audio key features and image key features of the video segment to be identified are spatiotemporally fused synchronously on the time axis to generate synchronous spatiotemporal features; Based on the data from the dissemination channels, the dissemination path of the video to be identified is traced and the spatiotemporal features are simultaneously authenticated to determine the video subject of the video to be identified; Based on the main body of the video, the video is identified as a violent and terrorist video, and when the main body of the video is a violent and terrorist video, the video segment to be identified is identified as a violent and terrorist video; The synchronous spatiotemporal feature authentication includes: Based on data from distribution channels, obtain related videos of the video segments to be identified on different platforms; The identifier axis is set according to the start and end image frames of the video segment to be identified, at a preset interval. Each associated video is authenticated using a corresponding identifier axis; Based on the corresponding authentication results, determine the set of target videos that may become the main subject of the associated videos.
2. The method for identifying violent and terrorist videos based on key feature fusion as described in claim 1, characterized in that, The key feature database includes an audio key feature database and a video key feature database; wherein... The audio key feature database includes: a scene audio key feature database, an event audio key feature database, and an object audio key feature database; among them, The scene audio key feature database is constructed based on the event audio key feature database and the object audio key feature database; The video key feature database includes a scene video key feature database, an event video key feature database, and an object video key feature database.
3. The method for identifying violent and terrorist videos based on key feature fusion as described in claim 1, characterized in that, The pre-establishment of a key feature database based on violent and terrorist videos includes the following steps: The design includes dynamic terrorist scenarios in two types: violent terrorist videos and violent terrorist audio. The simulated terrorist behavior of participants in a dynamic terrorist scenario was conducted through audio text simulation and video image representation simulation. Based on simulated terrorist behavior, key audio and video features of different participants were extracted. The key features of audio and video are further subdivided into three types of terrorist risk factors: scene, event, and object. The key features of video and audio under the corresponding terrorist risk factors are determined, thereby constructing a key feature database containing terrorist elements.
4. The method for identifying violent and terrorist videos based on key feature fusion as described in claim 1, characterized in that, The content features of the video segment to be identified are extracted by a content capture mechanism; wherein... The content capture mechanism is used to capture text, objects, people, and scene environment based on the frame images of the video segment to be identified; The content capture mechanism is also used to mark terrorist elements based on perceived terrorist risks after extracting content features, and to record time-varying information of these elements; among other things, Time-varying information and terrorist element markers are matched and aligned in time.
5. The method for identifying violent and terrorist videos based on key feature fusion as described in claim 1, characterized in that, The dissemination channel data includes the direct playback platforms and dissemination paths of the video segment to be identified; wherein, Direct playback platforms include the first platform, the second platform, and the third platform; among them, The first platform is a homogeneous content playback platform, used to determine the playback platform whose homogeneity with the video segment to be identified reaches a preset value; The second platform is a content playback platform, used to play the video segments that are identified as the parts of the video to be identified; The third platform is the secondary video playback platform, used to determine the playback platform that includes all content of the video segment to be identified; The propagation path includes the main propagation path and propagation branch paths; The main propagation path is a sequential path composed of the sequential propagation nodes of the video segment to be identified; The propagation branch path is the propagation path associated with each node in the sequential propagation nodes.
6. The method for identifying violent and terrorist videos based on key feature fusion as described in claim 1, characterized in that, The synchronous spatiotemporal fusion includes: Construct a timeline based on the video segment to be identified, and record frame images and audio data at each moment through the timeline; The frame image and audio data are detected by the homomorphic feature detection model configured in the key feature database to determine whether homomorphic features exist. When homomorphic features exist, they are used as key features of the target. Based on the key features of the target, feature sorting is performed on a time axis to generate a feature sequence; Based on the feature sequence, determine the sequence bits that simultaneously possess both key image features and key audio features at the same time, and form sequence pairs; Based on the sequence pairs, spatiotemporal weighted fusion is performed to generate synchronous spatiotemporal features.
7. The method for identifying violent and terrorist videos based on key feature fusion as described in claim 1, characterized in that, The tracing of the propagation path includes: Receive data on the distribution channels of the video segment to be identified; By comparing the data from the dissemination channels with the preset sample data, a source tracing operation is performed to determine the video metadata; Based on the video metadata, all publishing platforms of the video segment to be identified are determined, and the video subject of the earliest video segment to be identified among all publishing platforms is determined, thereby determining the platform from which the corresponding video was disseminated.
8. The method for identifying violent and terrorist videos based on key feature fusion as described in claim 1, characterized in that, The method further includes: Obtain multiple video frames from the video segment to be identified; Determine the terrorism parameters corresponding to each of multiple video frames; among them, Terrorism parameters are determined through key feature data; The terrorism parameter is used to identify the weight of each video frame in the video segment to be identified; Multiple video frames are fused based on their respective terrorism parameters to obtain a fused frame that associates the video segment to be identified with terrorism.
9. The method for identifying violent and terrorist videos based on key feature fusion as described in claim 1, characterized in that, The method further includes: Obtain assessment data for violent and terrorist videos; Based on the assessment data, hazard indicators and potential accident indicators are set for different key features on the video segments to be identified. Based on probability weights, calculate the first weight of each indicator in the terrorism hazard index and event probability index; Based on the scope weight, calculate the second weight of each indicator in the terrorist harm index and the event probability index; The final weights are averaged based on the first and second weights. The severity of the terrorist incident in the video segment to be identified is assessed based on the final weight.