Emergency detection method and device, equipment and storage medium
By using pre-trained event extraction models and risk prediction models to extract event types and elements from media data, and combining them with semantic modeling and clustering of event coding models, the problem of poor semantic expression ability in existing technologies is solved, and the efficiency and accuracy of emergency event detection are improved.
Patent Information
- Application Number
- CN202410474616.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-18
- Publication Date
- 2025-10-28
AI Technical Summary
In existing emergency event detection methods, the semantic expression ability of keyword matching is poor, and media data with similar semantics but no keyword hits are easily missed, resulting in low efficiency and accuracy of event detection.
By using a pre-trained event extraction model to extract event types and elements from media data, an event coding model to model text semantics, output word vectors of word elements, and perform clustering based on word vectors, combined with a risk prediction model to determine the risk probability of events, thus achieving event detection.
The efficiency and accuracy of emergency event detection are improved, and the semantic characterization capability of the event coding model is used to accurately extract event types and elements, thereby improving the event clustering effect.
Smart Images

Figure CN120849690A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technology, and in particular to a method, apparatus, device and storage medium for detecting emergencies. Background Technology
[0002] With the development of internet technology, a massive amount of multimedia data is generated online every day. Some of this information spreads and evolves extremely easily online, creating negative public opinion and generating adverse effects. Therefore, proactively detecting emerging events is particularly important. Event detection (ED) is a key task in information extraction, aiming to identify event trigger words and classify them into predefined event types.
[0003] In one existing event detection method, based on a pre-set keyword library, media data that matches the keywords published on social media is matched, a set of phrases is extracted from the media data, and then the phrases in the phrase set are clustered to obtain multiple phrase clusters. Each phrase cluster is considered as a sudden event. Then, a summary is extracted from the phrase set in a phrase cluster based on a pre-trained language model. The extracted summary is used as the event summary of the sudden event corresponding to the phrase cluster, and the event detection result is obtained.
[0004] Among the methods mentioned above, the keyword matching method has poor semantic expression capabilities and is prone to missing media data that are semantically similar but do not match the keywords. Moreover, the phrase clustering and event clustering methods are poor, resulting in low efficiency and accuracy in detecting sudden events. Summary of the Invention
[0005] This application provides a method, apparatus, device, and storage medium for detecting emergencies, which can improve the efficiency and accuracy of emergency detection.
[0006] In a first aspect, embodiments of this application provide a method for detecting sudden events, including:
[0007] For each piece of media data generated within the target time period, the event type and event elements of the media data are determined based on the corresponding text data and the trained event extraction model.
[0008] The text data corresponding to the media data, the title of the text data, the event type and event elements of the media data are used as inputs to a trained event encoding model, and the word vectors of the word elements of the text data are output. The word elements of the text data are the representation of the sequence corresponding to the text data.
[0009] Based on the word vectors of the word elements of the text data corresponding to each media data, the text data corresponding to the media data generated within the target time period are clustered to obtain at least one set of text data.
[0010] A text data set is defined as an event. For each text data set, the risk probability of the event corresponding to the text data set is determined according to the trained risk prediction model, and the detection result of the event is determined based on the risk probability of the event.
[0011] Secondly, embodiments of this application provide a device for detecting sudden events, comprising:
[0012] The determination module is used to determine the event type and event elements of each media data generated within a target time period, based on the text data corresponding to the media data and the trained event extraction model.
[0013] The first processing module is used to take the text data corresponding to the media data, the title of the text data, the event type and event elements of the media data as input to a trained event encoding model, and output the word vectors of the word elements of the text data, wherein the word elements of the text data are the representation of the sequence corresponding to the text data.
[0014] The clustering module is used to cluster the text data corresponding to the media data generated within the target time period based on the word vectors of the word elements of the text data corresponding to each media data, so as to obtain at least one set of text data.
[0015] The second processing module is used to identify a text data set as an event, and for each text data set, to determine the risk probability of the event corresponding to the text data set according to the trained risk prediction model, and to determine the detection result of the event based on the risk probability of the event.
[0016] In one embodiment, the determining module is used to:
[0017] Obtain the text data corresponding to the media data;
[0018] The event extraction model takes the text data as input and outputs the event type and event element of the media data.
[0019] In one embodiment, the event extraction model includes an information extraction sub-model, a first fully connected layer, a second fully connected layer, and a bilinear layer. The determination module is specifically used for:
[0020] The text data is segmented to obtain multiple words;
[0021] Determine the word elements of the text data, input the word elements and the plurality of words into the information extraction sub-model, and output the word vectors of the word elements and the word vectors of each word in the plurality of words;
[0022] The word vectors of the word elements are input into the first fully connected layer, and the event type of the media data is output. The word vectors of each word are input into the second fully connected layer, and the category of each word is output.
[0023] The word vector and category of each word are input into the bilinear layer and the conditional random field model, and the event elements of the media data are output.
[0024] In one embodiment, the clustering module is used for:
[0025] Select the first text data and the second text data from the text data corresponding to the media data generated within the target time period;
[0026] Calculate the similarity between the word vectors of the word elements in the first text data and the word vectors of the word elements in the second text data. If the similarity is greater than a preset threshold, then cluster the first text data and the second text data into one class to obtain a first text data set.
[0027] Continue clustering until the text data corresponding to the media data generated within the target time period is clustered.
[0028] In one embodiment, the clustering module is used for:
[0029] The text data corresponding to the media data generated within the target time period is assigned to N preset clustering nodes, so that each clustering node clusters the assigned text data and sends the obtained clustering results to the main clustering node. The main clustering node summarizes and clusters the received N clustering results to obtain the at least one set of text data. Each clustering node clusters the text data based on the similarity between the word vectors of the word elements of the text data corresponding to the two media data.
[0030] In one embodiment, the second processing module is used to:
[0031] For each piece of text data in the text dataset, the risk probability of the event corresponding to the text data is determined according to the risk prediction model;
[0032] The average risk probability of the event corresponding to the text data in the text dataset is determined as the risk probability of the event corresponding to the text dataset.
[0033] In one embodiment, the second processing module is used to:
[0034] The risk prediction model takes the event elements of the media data, the feature information of the publishing object of the media data, the behavioral feature information of the publishing object, and the text data as inputs, and outputs the risk probability of the event corresponding to the media data.
[0035] In one embodiment, the risk prediction model includes a first feature model, a second feature model, and a feature fusion model, and the second processing module is specifically used for:
[0036] The event elements of the media data, the feature information of the publishing object of the media data, and the behavioral feature information of the publishing object are input into the first feature model, and the encoded vector of the media data is output.
[0037] The event elements of the media data and the text data are input into the second feature model, and the word vectors of the word elements of the text data corresponding to the media data are output.
[0038] The encoding vector of the media data and the word vectors of the word elements of the corresponding text data are input into the feature fusion model, and the risk probability of the event corresponding to the media data is output.
[0039] In one embodiment, the second processing module is used to:
[0040] Obtain the feature information of the publishing object and the behavioral feature information of the publishing object of the media data.
[0041] In one embodiment, the first feature model includes multiple fully connected layers;
[0042] The first feature model is used to encode the event elements of the media data, the feature information of the publishing object of the media data, and the behavioral feature information of the publishing object to obtain the encoding vector of the media data.
[0043] In one embodiment, the second feature model includes multiple transformer models;
[0044] The feature fusion model is used to perform deep feature fusion on the encoding vector of the media data and the word vectors of the word elements of the text data corresponding to the media data to obtain the risk probability of the event corresponding to the media data.
[0045] Thirdly, embodiments of this application provide a computer device, including: a processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to perform the method of the first aspect.
[0046] Fourthly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on a computer program, cause the computer to perform the method as described in the first aspect.
[0047] Fifthly, embodiments of this application provide a computer program product containing instructions that, when executed on a computer, cause the computer to perform the method as described in the first aspect.
[0048] In summary, in this embodiment, for each piece of media data generated within a target time period, the event type and event elements of the media data are determined based on the corresponding text data and a trained event extraction model. Then, using the corresponding text data, the title of the text data, the event type, and the event elements as input to the trained event encoding model, word vectors of the text data's word elements are output. Based on the word vectors of the text elements corresponding to each piece of media data, the text data corresponding to the media data generated within the target time period are clustered to obtain at least one text data set. Finally, a text data set is defined as an event. For each text data set, the risk probability of the event corresponding to the text data set is determined based on a trained risk prediction model, and the event detection result is determined based on the event's risk probability. By modeling text semantics through an event coding model, the event coding model can accurately extract event types and event elements from media data, solving the problem of poor semantic expression capabilities in existing models. Then, by using the semantic characterization capabilities of the event extraction model, event coding is performed, outputting word vectors of word elements in the text data corresponding to the media data. Clustering is then performed based on the word vectors of word elements in the text data corresponding to the media data, improving the effect of event clustering, and thus improving the overall efficiency and accuracy of emergency event detection. Attached Figure Description
[0049] Figure 1 A schematic diagram illustrating an implementation scenario of an emergency event detection method provided in this application embodiment;
[0050] Figure 2 A flowchart of a method for detecting sudden events provided in an embodiment of this application;
[0051] Figure 3 This is a schematic diagram of the structure of an event extraction model provided in an embodiment of this application;
[0052] Figure 4 This is a schematic diagram of the clustering process of an event coding model provided in an embodiment of this application;
[0053] Figure 5 A schematic diagram of a clustering process provided in an embodiment of this application;
[0054] Figure 6 A schematic diagram illustrating the processing steps of a risk prediction model provided in this application embodiment;
[0055] Figure 7 This is a schematic diagram of the structure of an emergency detection device provided in an embodiment of this application;
[0056] Figure 8 This is a schematic block diagram of a computer device provided in an embodiment of this application. Detailed Implementation
[0057] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the embodiments of this application.
[0058] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of the embodiments of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the present application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.
[0059] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0060] Before introducing the technical solutions of the embodiments of this application, the relevant knowledge of the embodiments of this application will be introduced below:
[0061] 1. Artificial Intelligence (AI): This refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have the functions of perception, reasoning, and decision-making. AI technology is a comprehensive discipline involving a wide range of fields, including both hardware and software technologies. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operating / interactive systems, and mechatronics. AI software technologies mainly include computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0062] 2. Natural Language Processing (NLP): This is an important area within computer science and artificial intelligence. Research in this field involves natural language, the language people use in daily life, thus it is closely related to linguistics. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.
[0063] 3. Machine Learning (ML): This is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0064] 4. Deep Learning (DL): A branch of machine learning, it's an algorithm that attempts to perform high-level abstraction of data using multiple processing layers containing complex structures or multiple nonlinear transformations. Deep learning learns the inherent patterns and hierarchical representations of training sample data. The information gained during this learning process greatly aids in interpreting data such as text, images, and sound. The ultimate goal of deep learning is to enable machines to possess analytical and learning capabilities like humans, capable of recognizing data such as text, images, and sound. Deep learning is a complex machine learning algorithm, and its performance in speech and image recognition far surpasses previous related technologies.
[0065] 5. Neural Network (NN): A deep learning model in the fields of machine learning and cognitive science that mimics the structure and function of biological neural networks.
[0066] 6. Pre-trained models (PTMs), also known as foundational models or large models, refer to deep neural networks (DNNs) with a large number of parameters. These DNNs are trained on massive amounts of unlabeled data. Leveraging the function approximation capabilities of large-parameter DNNs, PTMs extract common features from the data. Through fine-tuning, efficient parameter fine-tuning (PEFT), and prompt-tuning techniques, they are suitable for downstream tasks. Therefore, pre-trained models can achieve ideal results in small-shot or zero-shot scenarios. PTMs can be categorized according to the data modality they process, such as language models (ELMO, BERT, GPT), visual models (swin-transformer, ViT, V-MOE), speech models (VALL-E), and multimodal models (ViBERT, CLIP, Flamingo, Gato). Multimodal models refer to models that establish feature representations for two or more data modalities. Pre-trained models are important tools for outputting AI-generated content (AIGC) and can also serve as a general interface connecting multiple specific task models.
[0067] The technical solutions provided in this application mainly involve technologies such as natural language processing, machine learning, and deep learning in artificial intelligence. The event extraction model, event encoding model, and risk prediction model in this application can be the language model or multimodal model in the above-mentioned pre-trained models, which can be specifically described through the following embodiments.
[0068] In related technologies, when detecting emergencies, the keyword matching method has poor semantic expression capabilities and is prone to missing media data that are semantically similar but do not match the keywords. Moreover, the phrase clustering method has poor event clustering effect, resulting in low efficiency and accuracy of emergency detection.
[0069] To address this issue, this application embodiment uses a pre-trained event extraction model to extract event types and event elements from each piece of media data generated within a target time period. Then, using the media data, the titles of the corresponding text data, the event types, and the event elements as input to a trained event encoding model, the event extraction model models text semantics and outputs word vectors for the word elements of the corresponding text data. Next, based on the word vectors of each piece of media data, the corresponding text data generated within the target time period is clustered to obtain at least one text data set. Each text data set is then identified as an event. Finally, a trained risk prediction model determines the risk probability of the event corresponding to each text data set, and the event detection result is determined based on the event's risk probability. By modeling text semantics through an event encoding model, the event encoding model can accurately extract event types and event elements from media data, solving the problem of poor semantic expression capabilities in existing models. Furthermore, the semantic characterization capability of the event extraction model is used for event encoding, outputting word vectors for the word elements of the corresponding text data. Clustering based on the word vectors of the corresponding text data improves the event clustering effect, thereby increasing the efficiency and accuracy of emergency event detection.
[0070] The embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving.
[0071] This application's embodiments can be applied to various scenarios, including but not limited to social networks, social media platforms, and content marketing platforms. In one embodiment, the method provided by this application can detect sudden events in content and information on social media platforms or content marketing platforms, effectively controlling their continued spread. For example, various types of information with negative public opinion orientation are generated daily on social media platforms, which, once spread, will have a significant negative impact on society. Therefore, it is necessary to detect events with negative public opinion orientation by real-time monitoring of various media data (text, images, and videos, etc.) on the platform. Specifically, firstly, structured information (including event type and event elements) is extracted from various media data using event extraction methods, and then the structured information is used to cluster the media data, aggregating media data with similar content together to form event-dimensional data. Finally, by combining the risk probability of the event and its content information, the likelihood of the event being negative public opinion information is determined, thereby enabling public opinion control.
[0072] It should be noted that the application scenarios described above are for illustrative purposes only and are not intended to limit the scope of this application. In specific implementations, the technical solutions provided in the embodiments of this application can be flexibly applied according to actual needs.
[0073] For example, Figure 1 This is a schematic diagram illustrating an implementation scenario of a sudden event detection method provided in an embodiment of this application, such as... Figure 1 As shown, the implementation scenario of this application involves server 1 and terminal device 2. Terminal device 2 can communicate with server 1 via a communication network. The communication network can be an intranet, the Internet, Global System for Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, voice communication network, or other wireless or wired networks.
[0074] In some possible implementations, terminal device 2 refers to a device with rich human-computer interaction methods, internet access capabilities, typically running various operating systems, and possessing strong processing power. Terminal devices can be smartphones, tablets, laptops, desktop computers, or smartwatches, but are not limited to these. Optionally, in this embodiment, terminal device 2 is equipped with various applications, such as video applications and news applications.
[0075] In some possible implementations, the terminal device 2 includes, but is not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc.
[0076] Figure 1 Server 1 in this context can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services. This application embodiment does not impose any limitations on this. In this application embodiment, server 1 can be the backend server of an application installed on terminal device 2.
[0077] In some possible ways, Figure 1 An exemplary terminal device and a server are shown, but in practice, other numbers of terminal devices and servers may be included, and this application does not limit this.
[0078] In some embodiments, when it is necessary to detect sudden events, server 1 can use the method provided in the embodiments of this application to first train the event extraction model, the event encoding model, and the risk prediction model to obtain the trained event extraction model, the trained event encoding model, and the trained risk prediction model. Then, specifically when detecting sudden events, terminal device 2 can send media data generated within the target time period to the server. After receiving the media data generated within the target time period, the server executes the sudden event detection method provided in the embodiments of this application to obtain the detection results of one or more events generated within the target time period, and thus obtain the possibility that the one or more events are negative public opinion information.
[0079] It is understood that in the specific implementation of this application, data related to user information (such as the characteristic information and behavioral characteristic information of the publishing object, when the publishing object is a user) are involved. When the method of this application embodiment is applied to a specific product or technology, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0080] The technical solutions of the embodiments of this application will be described in detail below:
[0081] Figure 2 This is a flowchart illustrating a method for detecting sudden events provided in an embodiment of this application. The executing entity of this embodiment is a device with sudden event detection functionality, such as a server. Figure 2 As shown, the method may include:
[0082] S101. For each piece of media data generated within the target time period, determine the event type and event elements of the media data based on the corresponding text data and the trained event extraction model.
[0083] Specifically, in one embodiment, the terminal device can collect media data within a preset period, such as a target time. The terminal device periodically collects media data within the target time period and sends the media data within the target time period to the server, which then executes S101-S104.
[0084] In another embodiment, the server may collect media data within a preset period, such as a target time period. The server periodically collects media data within the target time period and then executes S101-S104.
[0085] Specifically, media data refers to various forms of media content, which can be text, images, audio, or video. The corresponding text data refers to the text information related to this media content. For example, a video might have a title, description, or subtitles; these are all text data related to the video. If the media data is text, the corresponding text data is the text itself. If the media data is an image, the text corresponding to the image needs to be extracted as the corresponding text data. If the media data is video, the text corresponding to the video needs to be extracted as the corresponding text data. For example, if 1000 pieces of media data are generated within a target time period, for each of these 1000 pieces of media data, the event type and event elements of the media data are determined based on the corresponding text data and the trained event extraction model.
[0086] In this embodiment, the event extraction model, the event encoding model, and the risk prediction model mentioned below are all pre-trained network models. In one embodiment, the event extraction model, the event encoding model, and the risk prediction model mentioned below can all be large models, such as language models or multimodal models mentioned above.
[0087] Specifically, in one feasible approach, in step S101, the event type and event elements of the media data are determined based on the text data corresponding to the media data and the trained event extraction model. Specifically, this can be achieved by:
[0088] S1011. Obtain the text data corresponding to the media data.
[0089] The media data can be text, images, or videos. If the media data is text, the corresponding text data is a segment of text itself, which is the text data. If the media data is an image, the text corresponding to the image can be obtained through image recognition, which is the text data corresponding to the image. If the media data is a video, the text description of the video can be used as the text data corresponding to the video, or the text data corresponding to the video can be obtained through other means. This embodiment does not impose any restrictions on this.
[0090] S1012. Using the text data corresponding to the media data as input to the event extraction model, output the event type and event element of the media data.
[0091] After acquiring the text data corresponding to the media data, the text data for each data point is input into the event extraction model, which outputs the event type and event elements of the media data. The event type can be predefined according to business needs; for example, the event type can be a meeting, negotiation, inauguration (taking office), dismissal (stepping down), removal from office, or expulsion. Event elements can include event time, event location, subject, and object. In other scenarios, the event type and event elements can be defined differently, and this embodiment does not impose any restrictions. In this embodiment, the event extraction model possesses basic Chinese NLP capabilities, such as lexical analysis, syntactic analysis, and semantic parsing.
[0092] Optionally, in one implementable approach, the event extraction model includes an information extraction sub-model, a first fully connected layer, a second fully connected layer, a bilinear layer, and a conditional random field model. The text data corresponding to the media data serves as the input to the event extraction model, and the output is the event type and event elements of the media data. Specifically, it can be:
[0093] S1. Perform word segmentation on the text data corresponding to the media data to obtain multiple words.
[0094] Specifically, word segmentation is the process of dividing a continuous text sequence into meaningful lexical units. The text data corresponding to the media data is a segment of text, which is then segmented into words. For example, the segment can be segmented based on a set number of words (e.g., 2) to obtain multiple words. During the word segmentation process, word boundaries can also be determined based on vocabulary, grammatical rules, or statistical information from a dictionary. There are various word segmentation methods, and this embodiment does not limit them.
[0095] S2. Determine the word elements of the text data corresponding to the media data, extract the word elements and multiple words into the sub-model, and output the word vector of the word elements and the word vector of each word in the multiple words.
[0096] Specifically, the classification element (CLS) of text data is the representation of the corresponding sequence of text data. This classification element is used to indicate the beginning of the text data sequence. The first position of the text data sequence input to the information extraction sub-model is marked as the CLS, which is used to represent the summary information of the entire sequence. The information extraction sub-model learns to use the representation of the CLS position to perform various classification tasks, such as text classification and sentiment analysis. The role of CLS is to enable the model to capture the global information of the entire input sequence (i.e., the text data sequence) and perform subsequent classification or prediction tasks based on this global information.
[0097] S3. Input the word vectors of the word elements into the first fully connected layer, and output the event type of the media data. Input the word vector of each word into the second fully connected layer, and output the category of each word.
[0098] S4. Input the category of each word into the bilinear layer and conditional random field model, and output the event elements of the media data.
[0099] Figure 3 This is a schematic diagram of the structure of an event extraction model provided in an embodiment of this application, such as... Figure 3 As shown, the event extraction model in this embodiment includes an information extraction sub-model 10, a first fully connected layer 20, a second fully connected layer 30, a bilinear layer 40, and a conditional random field model 50. The information extraction sub-model 10 can be an LTP model. The LTP model serves as a hot-start model, enabling the event extraction model to possess basic NLP capabilities, such as lexical analysis, syntax parsing, and semantic parsing. Combined with... Figure 3 As shown, the text data corresponding to the media data is used as the input of the event extraction model, and the output is the event type and event element of the media data. Specifically, the text data corresponding to the media data is first segmented into words to obtain multiple words such as S1, S2, ..., S8, and the word element CLS of the text data corresponding to the media data is determined. The word element CLS and the 8 words (S1-S8) are input into the information extraction sub-model, and the word vector of the word element (CLS-E) and the word vector of each of the 8 words (E1-E8) are output. Next, the word vector CLS-E of the word element is input into the first fully connected layer 20, and the event type of the media data is output. The word vector (E1-E8) of each word is input into the second fully connected layer 30, and the category of each word is output. Finally, the categories of each word are input into the bilinear layer 40 and the conditional random field model 50, and the event elements (L1-L8) of the media data are output.
[0100] Optionally, in one embodiment, when training the event extraction model, a loss function is constructed based on the predicted event elements, event types, and event types of the sample data output by the conditional random field model. The model parameters of the event extraction model are adjusted using the loss function until the training stop condition is met, resulting in a trained event extraction model. Thus, the trained event extraction model possesses the function of extracting the event types and event elements of the media data from the text data corresponding to the media data.
[0101] S102. The text data corresponding to the media data, the title of the text data, the event type and event element of the media data are used as inputs to the trained event encoding model, and the output is the word vector of the word element of the text data. The word element of the text data is the representation of the sequence corresponding to the text data.
[0102] Specifically, the event encoding model can be a Transformer model, which is a neural network model based on a self-attention mechanism. The Transformer model can be used for tasks such as text classification, machine translation, named entity recognition, and sentiment analysis. It uses a self-attention mechanism to globally model each element in the sequence and establish connections between elements, thereby effectively capturing information at any position in the sequence.
[0103] In this embodiment, by pre-training an event encoding model, the event encoding model can output word vectors (CLS) of the word elements of the input media data based on the text data corresponding to the input media data, the title of the text data, the event type of the media data, and the event elements of the text data. The word elements of the text data represent the sequence corresponding to the text data. The event encoding model uses the event elements of the media data as core keywords, and adds the corresponding text data and the title of the text data to enrich the semantic information.
[0104] S103. Based on the word vectors of the word elements of the text data corresponding to each media data, cluster the text data corresponding to the media data generated within the target time period to obtain at least one text data set.
[0105] Optionally, in one implementable manner, in S103, based on the word vectors of the word elements of the text data corresponding to each media data, the text data corresponding to the media data generated within the target time period are clustered to obtain at least one text data set, which can specifically be:
[0106] S1031. Select the first text data and the second text data from the text data corresponding to the media data generated within the target time period.
[0107] S1032. Calculate the similarity between the word vectors of the word elements of the first text data and the word vectors of the word elements of the second text data. If the similarity is greater than a preset threshold, then cluster the first text data and the second text data into one class to obtain the first text data set.
[0108] S1033. Continue clustering until the text data corresponding to the media data generated within the target time period is clustered.
[0109] Specifically, in one embodiment, two event coding models with the same structure can be pre-trained. Based on the outputs of the two event coding models with the same structure, the text data corresponding to the two sample media data are clustered. Specifically, based on the similarity between the word vectors of the word elements of the first text data and the word vectors of the word elements of the second text data output by the two event coding models with the same structure, it is determined whether the first text data and the second text data are the same event. If so, the first text data and the second text data are combined into a text data set.
[0110] Figure 4 This is a schematic diagram of the clustering process of an event coding model provided in an embodiment of this application, as shown below. Figure 4 As shown, there are two event encoding models with the same structure: Event Encoding Model 1 and Event Encoding Model 2. The input of Event Encoding Model 1 is the text data corresponding to the first media data, the title of the text data, the event type of the first media data, and the event element. The output of Event Encoding Model 1 is the word vector of the word element of the text data, i.e., CLS word vector 1. The input of Event Encoding Model 2 is the text data corresponding to the second media data, the title of the text data, the event type of the second media data, and the event element. The output of Event Encoding Model 2 is the word vector of the word element of the text data, i.e., CLS word vector 2. Next, CLS word vector 1 and CLS word vector 2 are input into the feature interaction layer. The feature interaction layer calculates the similarity between CLS word vector 1 and CLS word vector 2. If the similarity is greater than a preset threshold, it is determined that the event corresponding to the first media data and the event corresponding to the second media data are the same event. The text data corresponding to the first media data and the text data corresponding to the second media data are clustered into one class to obtain the first text data set.
[0111] The above method involves sequentially clustering the text data corresponding to the media data generated within the target time period. Further, to improve the throughput of the event encoding model, distributed streaming clustering can be performed, deploying multiple clustering nodes. Each clustering node clusters the received text data and sends the clustering results to the master clustering node. The master clustering node then aggregates and clusters the received clustering results to obtain at least one set of text data. This will be described in detail below via S1031'.
[0112] In another feasible approach, in step S103, based on the word vectors of the word elements of the text data corresponding to each media data segment, the text data corresponding to the media data generated within the target time period are clustered to obtain at least one set of text data, which can specifically be:
[0113] S1031' Distribute the text data corresponding to the media data generated within the target time period to N preset clustering nodes, so that each clustering node clusters the distributed text data and sends the obtained clustering results to the main clustering node, which then summarizes and clusters the received N clustering results to obtain at least one set of text data. Each clustering node clusters the text data corresponding to two media data based on the similarity between the word vectors of the word elements.
[0114] Specifically, Figure 5 A schematic diagram of a clustering process provided in an embodiment of this application, such as... Figure 5 As shown, this embodiment can use the single-pass algorithm for distributed streaming clustering, deploying N clustering nodes. The text data corresponding to the media data generated within the target time period is distributed to the N clustering nodes, for example, by equal distribution. Each clustering node clusters the received text data. Each clustering node can cluster based on the similarity between the word vectors of the word elements of the text data corresponding to two media data. Each clustering node sends the obtained clustering results to the master clustering node. The master clustering node summarizes and clusters the received N clustering results to obtain at least one set of text data.
[0115] S104. Define a text data set as an event. For each text data set, determine the risk probability of the event corresponding to the text data set based on the trained risk prediction model, and determine the detection result of the event based on the risk probability of the event.
[0116] Specifically, after obtaining at least one text dataset, each text dataset can be identified as an event, and the detection result for each event can be determined. For example, if four text datasets are obtained, the detection results for the events corresponding to each of the four text datasets can be determined separately. For each text dataset, the risk probability of the event corresponding to that text dataset is determined based on a pre-trained risk prediction model. The risk prediction model is pre-trained and can predict the risk probability of each text dataset.
[0117] Optionally, in one implementable approach, in step S104, the risk probability of the event corresponding to the text dataset is determined based on the trained risk prediction model. Specifically, this can be:
[0118] S1041. For each text data in the text data set, determine the risk probability of the event corresponding to the text data according to the risk prediction model.
[0119] S1042. The average risk probability of the event corresponding to the text data in the text data set is determined as the risk probability of the event corresponding to the text data set.
[0120] Optionally, in one implementable manner, S1041 determines the risk probability of the event corresponding to the text data based on a risk prediction model, specifically as follows:
[0121] S10411. The risk prediction model takes the event elements of the media data, the feature information of the media data's publishing object, the behavioral feature information of the publishing object, and the text data corresponding to the media data as inputs and outputs the risk probability of the event corresponding to the media data.
[0122] Specifically, the recipients of media data are the publishers of the media data. These recipients can encompass multiple fields and groups, such as the general public, businesses and merchants, government agencies, researchers and scholars, or other media organizations. The characteristics of the recipients of media data can include, for example, profile features such as age, gender, and interests; behavioral characteristics such as information acquisition behavior, interaction behavior, and time-related features; information acquisition behavior such as historical browsing and click data; interaction behavior such as likes, comments, and shares; and time-related features such as periods of high activity levels for the recipients.
[0123] Optionally, in one implementable manner, the risk prediction model includes a first feature model, a second feature model, and a feature fusion model, wherein S10411 above can specifically be:
[0124] S11. Input the event elements of the media data, the feature information of the media data publishing object, and the behavioral feature information of the publishing object into the first feature model, and output the encoding vector of the media data.
[0125] S12. Input the event elements of the media data and the corresponding text data of the media data into the second feature model, and output the word vectors of the word elements of the corresponding text data of the media data.
[0126] S13. Input the encoding vector of the media data and the word vector of the word element of the text data corresponding to the media data into the feature fusion model, and output the risk probability of the event corresponding to the media data.
[0127] Specifically, the feature fusion model can be the BERT model.
[0128] Furthermore, the method in this embodiment may also include:
[0129] S14. Obtain the characteristic information and behavioral characteristic information of the media data publishing object.
[0130] Optionally, in one implementable manner, the first feature model includes multiple fully connected layers;
[0131] The first feature model is used to encode the event elements of the media data, the feature information of the media data publishing object, and the behavioral feature information of the publishing object, to obtain the encoding vector of the media data.
[0132] Optionally, in one feasible approach, the second feature model includes multiple transformer models. The feature fusion model is used to perform deep feature fusion on the encoded vector of the media data and the word vectors of the word elements of the corresponding text data to obtain the risk probability of the event corresponding to the media data.
[0133] The following is combined with Figure 6 This paper presents the structure and processing procedure of a risk prediction model. Figure 6 This is a schematic diagram illustrating the processing steps of a risk prediction model provided in an embodiment of this application, as shown below. Figure 6 As shown, the risk prediction model includes a first feature model 100, a second feature model 200, and a feature fusion model 300. The first feature model 100 can be a multi-layer fully connected layer, and the second feature model 200 can include multiple transformer models.
[0134] The first feature model 100 takes as input event elements of the media data, feature information of the media data's publishing object, and behavioral feature information of the publishing object, and outputs as the encoded vector of the media data. The second feature model 200 takes as input event elements of the media data and the corresponding text data, and outputs as word vectors of the word elements of the corresponding text data. Then, the encoded vector of the media data and the word vectors of the corresponding text data are input into the feature fusion model, which outputs the risk probability of the event corresponding to the media data.
[0135] In one embodiment, the risk probability of an event can be preset with a first threshold and a second threshold, where the first threshold is greater than the second threshold. For example, if the risk probability of an event is greater than the first threshold, the event is determined to be an important event; if the risk probability of an event is less than the first threshold but greater than the second threshold, the event is determined to be a high-profile, highly contagious event; and if the risk probability of an event is less than the second threshold, the event is determined to be a recommended event. The risk probability of an event and the detection result of the event can also be defined according to business needs. This embodiment is only an example and does not impose any limitations on this.
[0136] In this embodiment of the application, the risk prediction model introduces the feature information and behavioral feature information of the publishing object to enrich the semantic structured information. From the model perspective, the fusion of semantic structured information and text information optimizes the performance of the risk prediction model in obtaining the risk probability of an event.
[0137] The emergency event detection method provided in this embodiment determines the event type and event elements of each media data point generated within a target time period, based on the corresponding text data and a trained event extraction model. Then, using the corresponding text data, title, event type, and event elements as input to a trained event encoding model, the model outputs word vectors for the word elements of the text data. Based on the word vectors of the text elements corresponding to each media data point, the text data corresponding to the media data generated within the target time period is clustered to obtain at least one text data set. Finally, a text data set is defined as an event. For each text data set, the risk probability of the event corresponding to the text data set is determined based on a trained risk prediction model, and the event detection result is determined based on the event's risk probability. By modeling text semantics through an event coding model, the event coding model can accurately extract event types and event elements from media data, solving the problem of poor semantic expression capabilities in existing models. Then, by using the semantic characterization capabilities of the event extraction model, event coding is performed, outputting word vectors of word elements in the text data corresponding to the media data. Clustering is then performed based on the word vectors of word elements in the text data corresponding to the media data, improving the effect of event clustering, and thus improving the overall efficiency and accuracy of emergency event detection.
[0138] Figure 7 This is a schematic diagram of the structure of a sudden event detection device provided in an embodiment of this application, as shown below. Figure 7 As shown, the device may include: a determination module 11, a first processing module 12, a clustering module 13, and a second processing module 14.
[0139] The determination module 11 is used to determine the event type and event elements of each media data generated within the target time period, based on the corresponding text data and the trained event extraction model.
[0140] The first processing module 12 is used to take the text data corresponding to the media data, the title of the text data, the event type and event element of the media data as input to the trained event encoding model, and output the word vectors of the word elements of the text data. The word elements of the text data are the representation of the sequence corresponding to the text data.
[0141] The clustering module 13 is used to cluster the text data corresponding to the media data generated within the target time period based on the word vectors of the word elements of the text data corresponding to each media data, so as to obtain at least one set of text data.
[0142] The second processing module 14 is used to identify a text data set as an event, and for each text data set, to determine the risk probability of the event corresponding to the text data set according to the trained risk prediction model, and to determine the detection result of the event based on the risk probability of the event.
[0143] In one embodiment, the determining module 11 is used to:
[0144] Retrieve the text data corresponding to the media data;
[0145] The event extraction model takes the text data corresponding to the media data as input and outputs the event type and event element of the media data.
[0146] In one embodiment, the event extraction model includes an information extraction sub-model, a first fully connected layer, a second fully connected layer, and a bilinear layer. The determination module 11 is specifically used for:
[0147] The text data corresponding to the media data is segmented into words to obtain multiple words;
[0148] Determine the word elements of the text data corresponding to the media data, extract the sub-model from the input information of the word elements and multiple words, and output the word vector of the word elements and the word vector of each word in the multiple words;
[0149] Input the word vectors of the word elements into the first fully connected layer, and output the event type of the media data. Input the word vectors of each word into the second fully connected layer, and output the category of each word.
[0150] Each word's category is input into a bilinear layer, which outputs event elements of the media data.
[0151] In one embodiment, the clustering module 13 is used to:
[0152] Select the first and second text data from the text data corresponding to the media data generated within the target time period;
[0153] Calculate the similarity between the word vectors of the word elements in the first text data and the word vectors of the word elements in the second text data. If the similarity is greater than a preset threshold, then cluster the first text data and the second text data into one class to obtain the first text data set.
[0154] Continue clustering until the text data corresponding to the media data generated within the target time period is clustered.
[0155] In one embodiment, the clustering module 13 is used to:
[0156] The text data corresponding to the media data generated within the target time period is assigned to N preset clustering nodes, so that each clustering node clusters the assigned text data and sends the clustering results to the main clustering node. The main clustering node summarizes and clusters the received N clustering results to obtain at least one set of text data. Each clustering node clusters the text data corresponding to two media data based on the similarity between the word vectors of the word elements.
[0157] In one embodiment, the second processing module 14 is used to:
[0158] For each piece of text data in the text dataset, the risk probability of the event corresponding to the text data is determined based on the risk prediction model.
[0159] The average risk probability of the event corresponding to the text data in the text dataset is determined as the risk probability of the event corresponding to the text dataset.
[0160] In one embodiment, the second processing module 14 is used to:
[0161] The risk prediction model takes the event elements of media data, the feature information of the media data's publisher, the behavioral feature information of the publisher, and text data as inputs and outputs the risk probability of the event corresponding to the media data.
[0162] In one embodiment, the risk prediction model includes a first feature model, a second feature model, and a feature fusion model, and the second processing module 14 is specifically used for:
[0163] Input the event elements of the media data, the feature information of the media data's publishing object, and the behavioral feature information of the publishing object into the first feature model, and output the encoded vector of the media data;
[0164] The event elements and text data of the media data are input into the second feature model, and the word vectors of the word elements of the text data corresponding to the media data are output.
[0165] The model inputs the encoded vector of the media data and the word vectors of the corresponding text data into the feature fusion model, and outputs the risk probability of the event corresponding to the media data.
[0166] In one embodiment, the second processing module 14 is used to:
[0167] Obtain the characteristic information and behavioral characteristic information of the media data publishing object.
[0168] In one embodiment, the first feature model includes multiple fully connected layers;
[0169] The first feature model is used to encode the event elements of the media data, the feature information of the media data publishing object, and the behavioral feature information of the publishing object, to obtain the encoding vector of the media data.
[0170] In one embodiment, the second feature model includes multiple transformer models;
[0171] The feature fusion model is used to perform deep feature fusion on the encoded vector of media data and the word vectors of the word elements of the corresponding text data to obtain the risk probability of the event corresponding to the media data.
[0172] It should be understood that the device embodiments and method embodiments can correspond to each other, and similar descriptions can be found in the method embodiments. To avoid repetition, further details are omitted here. Specifically, Figure 7 The emergency detection device shown can execute the method embodiment corresponding to the computer device, and the foregoing and other operations and / or functions of each module in the device are respectively for implementing the method embodiment corresponding to the computer device. For the sake of brevity, they will not be described in detail here.
[0173] The above description, in conjunction with the accompanying drawings, of the emergency detection device and information prediction device of this application embodiments from the perspective of functional modules. It should be understood that these functional modules can be implemented in hardware, in software instructions, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in this application can be completed by integrated logic circuits in the processor's hardware and / or by software instructions. The steps of the methods disclosed in the embodiments of this application can be directly manifested as execution by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. Optionally, the software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps in the above method embodiments.
[0174] Figure 8 This is a schematic block diagram of a computer device provided in an embodiment of this application.
[0175] like Figure 8 As shown, the computer device may include:
[0176] The system includes a memory 310 and a processor 320. The memory 310 stores computer programs and transfers the program code to the processor 320. In other words, the processor 320 can retrieve and run the computer program from the memory 310 to implement the methods described in the embodiments of this application.
[0177] For example, the processor 320 can be used to execute the above-described method embodiments according to instructions in the computer program.
[0178] In some embodiments of this application, the processor 320 may include, but is not limited to:
[0179] General-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0180] In some embodiments of this application, the memory 310 includes, but is not limited to:
[0181] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0182] In some embodiments of this application, the computer program may be divided into one or more modules, which are stored in the memory 310 and executed by the processor 320 to complete the method provided in this application. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.
[0183] like Figure 8 As shown, the computer device may further include:
[0184] Transceiver 330, which can be connected to processor 320 or memory 310.
[0185] The processor 320 can control the transceiver 330 to communicate with other devices; specifically, it can send information or data to other devices or receive information or data sent by other devices. The transceiver 330 may include a transmitter and a receiver. The transceiver 330 may further include antennas, and the number of antennas may be one or more.
[0186] It should be understood that the various components in the electronic device are connected through a bus system, which includes a data bus, a power bus, a control bus, and a status signal bus.
[0187] This application also provides a computer storage medium storing a computer program thereon, which, when executed by a computer, enables the computer to perform the methods of the above-described method embodiments. Alternatively, this application also provides a computer program product containing instructions that, when executed by a computer, cause the computer to perform the methods of the above-described method embodiments.
[0188] When implemented using software, it can be implemented entirely or partially as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0189] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the embodiments of this application.
[0190] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or modules may be electrical, mechanical, or other forms.
[0191] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. For example, the functional modules in the various embodiments of this application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0192] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the embodiments of this application should be included within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.
Claims
1. A method for detecting sudden events, characterized in that, include: For each piece of media data generated within the target time period, the event type and event elements of the media data are determined based on the corresponding text data and the trained event extraction model. The text data corresponding to the media data, the title of the text data, the event type and event elements of the media data are used as inputs to a trained event encoding model, and the word vectors of the word elements of the text data are output. The word elements of the text data are the representation of the sequence corresponding to the text data. Based on the word vectors of the word elements of the text data corresponding to each media data, the text data corresponding to the media data generated within the target time period are clustered to obtain at least one set of text data. A text data set is defined as an event. For each text data set, the risk probability of the event corresponding to the text data set is determined according to the trained risk prediction model, and the detection result of the event is determined based on the risk probability of the event.
2. The method according to claim 1, characterized in that, The step of determining the event type and event elements of the media data based on the text data corresponding to the media data and the trained event extraction model includes: Obtain the text data corresponding to the media data; The event extraction model takes the text data as input and outputs the event type and event element of the media data.
3. The method according to claim 1, characterized in that, The event extraction model includes an information extraction sub-model, a first fully connected layer, a second fully connected layer, a bilinear layer, and a conditional random field model. The model takes the text data as input and outputs the event types and event elements of the media data, including: The text data is segmented to obtain multiple words; Determine the word elements of the text data, input the word elements and the plurality of words into the information extraction sub-model, and output the word vectors of the word elements and the word vectors of each word in the plurality of words; The word vectors of the word elements are input into the first fully connected layer, and the event type of the media data is output. The word vectors of each word are input into the second fully connected layer, and the category of each word is output. The word vector and category of each word are input into the bilinear layer and the conditional random field model, and the event elements of the media data are output.
4. The method according to claim 1, characterized in that, The step involves clustering the text data corresponding to the media data generated within the target time period based on the word vectors of the word elements of the text data corresponding to each media data segment, to obtain at least one text data set, including: Select the first text data and the second text data from the text data corresponding to the media data generated within the target time period; Calculate the similarity between the word vectors of the word elements in the first text data and the word vectors of the word elements in the second text data. If the similarity is greater than a preset threshold, then cluster the first text data and the second text data into one class to obtain a first text data set. Continue clustering until the text data corresponding to the media data generated within the target time period is clustered.
5. The method according to claim 1, wherein The step involves clustering the text data corresponding to the media data generated within the target time period based on the word vectors of the word elements of the text data corresponding to each media data segment, to obtain at least one text data set, including: The text data corresponding to the media data generated within the target time period is assigned to N preset clustering nodes, so that each clustering node clusters the assigned text data and sends the obtained clustering results to the main clustering node. The main clustering node summarizes and clusters the received N clustering results to obtain the at least one set of text data. Each clustering node clusters the text data based on the similarity between the word vectors of the word elements of the text data corresponding to the two media data.
6. The method according to any one of claims 1-5, characterized in that, Determining the risk probability of the event corresponding to the text dataset based on the trained risk prediction model includes: For each piece of text data in the text dataset, the risk probability of the event corresponding to the text data is determined according to the risk prediction model; The average risk probability of the event corresponding to the text data in the text dataset is determined as the risk probability of the event corresponding to the text dataset.
7. The method according to claim 6, characterized in that, Determining the risk probability of the event corresponding to the text data based on the risk prediction model includes: The risk prediction model takes the event elements of the media data, the feature information of the publishing object of the media data, the behavioral feature information of the publishing object, and the text data as inputs, and outputs the risk probability of the event corresponding to the media data.
8. The method according to claim 7, characterized in that, The risk prediction model includes a first feature model, a second feature model, and a feature fusion model. The input to the risk prediction model consists of the event elements of the media data, the feature information of the media data's publishing object, the behavioral feature information of the publishing object, and the text data. The output includes the risk probability of the event corresponding to the media data. The event elements of the media data, the feature information of the publishing object of the media data, and the behavioral feature information of the publishing object are input into the first feature model, and the encoded vector of the media data is output. The event elements of the media data and the text data are input into the second feature model, and the word vectors of the word elements of the text data corresponding to the media data are output. The encoding vector of the media data and the word vectors of the word elements of the corresponding text data are input into the feature fusion model, and the risk probability of the event corresponding to the media data is output.
9. The method according to claim 8, characterized in that, The method further includes: Obtain the feature information of the publishing object and the behavioral feature information of the publishing object of the media data.
10. The method according to claim 8, characterized in that, The first feature model includes multiple fully connected layers; The first feature model is used to encode the event elements of the media data, the feature information of the publishing object of the media data, and the behavioral feature information of the publishing object to obtain the encoding vector of the media data.
11. The method according to claim 8, characterized in that, The second feature model includes multiple transformer models; The feature fusion model is used to perform deep feature fusion on the encoding vector of the media data and the word vectors of the word elements of the text data corresponding to the media data to obtain the risk probability of the event corresponding to the media data.
12. A device for detecting sudden events, characterized in that, include: The determination module is used to determine the event type and event elements of each media data generated within a target time period, based on the text data corresponding to the media data and the trained event extraction model. The first processing module is used to take the text data corresponding to the media data, the title of the text data, the event type and event elements of the media data as input to a trained event encoding model, and output the word vectors of the word elements of the text data, wherein the word elements of the text data are the representation of the sequence corresponding to the text data. The clustering module is used to cluster the text data corresponding to the media data generated within the target time period based on the word vectors of the word elements of the text data corresponding to each media data, so as to obtain at least one set of text data. The second processing module is used to identify a text data set as an event, and for each text data set, to determine the risk probability of the event corresponding to the text data set according to the trained risk prediction model, and to determine the detection result of the event based on the risk probability of the event.
13. A computer device, characterized in that, include: A processor and a memory, the memory being used to store a computer program, the processor being used to invoke and run the computer program stored in the memory to perform the method of any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, Includes instructions that, when executed on a computer program, cause the computer to perform the method as described in any one of claims 1 to 11.
15. A computer program product containing instructions, characterized in that, When the instructions are executed on a computer, the computer causes the computer to perform the method of any one of claims 1 to 11.