An event recognition method, device, electronic device and storage medium
Through the combination of clustering algorithm and large language model, the problem of low efficiency of manually identifying events is solved, and the effect of quickly filtering out high-frequency event information is achieved.
Patent Information
- Application Number
- CN202510031201.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-01-08
AI Technical Summary
In the prior art, the efficiency of manually identifying events to be processed is low, especially when filtering out effective event information from massive information, it is insufficient.
Clustering algorithm is used to determine the categories of representation events from the text, and a large language model is used to generate topic perspectives, and the text of pending events is determined through similarity matching.
It improves the efficiency of identifying events to be processed and can quickly filter out event information with high frequency representation from a large amount of information.
Smart Images

Figure CN119441497B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an event recognition method, apparatus, electronic device, and storage medium. Background Art
[0002] With the development of network technology, the public can upload opinions and viewpoints on specified events on relevant platforms through the network. For example, relevant departments can establish a problem feedback platform for the public. Correspondingly, the public can upload the information of the problems to be reflected to this problem feedback platform. The information of the problems to be reflected may be an event occurring in a certain area. However, the information received by the platform may be massive, and there may be a large amount of information on repeated events among the received information. Each piece of information may also contain invalid content (such as garbled characters) that has nothing to do with the event that the user needs to reflect. In one way, the received texts can be statistically analyzed one by one manually, and the text representing the event to be processed can be determined.
[0003] However, the efficiency of manually identifying the event to be processed is low. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide an event recognition method, apparatus, electronic device, and storage medium to improve the efficiency of identifying the event to be processed. The specific technical solutions are as follows:
[0005] In the first aspect of the embodiments of the present application, first, an event recognition method is provided. The method includes:
[0006] Obtain multiple texts containing event information of events as the first texts;
[0007] Based on a preset clustering algorithm, determine texts in which the represented events belong to the category to be processed from each of the first texts as the second texts;
[0008] Input each of the second texts and a first prompt word into a large language model to obtain a first topic view of the texts of the category to be processed; wherein, the first prompt word is used to instruct the large language model to generate the topic view of the input text;
[0009] Use the large language model to determine texts matching the first topic view from each of the second texts as the third texts representing the event to be processed.
[0010] In some embodiments, the step of determining texts in which the represented events belong to the category to be processed from each of the first texts as the second texts based on a preset clustering algorithm includes:
[0011] Encode each of the first texts to obtain a first embedding vector of each of the first texts;
[0012] Cluster the first embedding vectors based on a preset clustering algorithm to obtain at least one vector cluster;
[0013] Select one vector cluster belonging to the category to be processed from the obtained at least one vector cluster;
[0014] Use the first text represented by the first embedding vectors included in the selected vector cluster as the second text.
[0015] In some embodiments, before encoding each first text to obtain the first embedding vector of each first text, the method further includes:
[0016] Input the average length, standard deviation of each first text, and a third prompt word into a large language model to obtain a first parameter for dimensionality reduction processing and a second parameter for clustering;
[0017] Wherein, the first parameter includes at least one of the following: the dimension of the embedding vector after dimensionality reduction, the minimum distance between each embedding vector after dimensionality reduction, and the number of adjacent points when constructing a neighborhood after dimensionality reduction; the second parameter includes at least one of the following: the minimum number of vectors included in a vector cluster, the minimum number of vectors in the neighborhood of the center of a vector cluster; the third prompt word is used to: instruct the large language model to generate parameters for dimensionality reduction processing and parameters for clustering based on the average length and standard deviation of the input text;
[0018] The encoding of each first text to obtain the first embedding vector of each first text includes:
[0019] Encode each first text and perform dimensionality reduction processing on the encoding result according to the first parameter to obtain the first embedding vector of each first text;
[0020] The clustering of the first embedding vectors based on a preset clustering algorithm to obtain at least one vector cluster includes:
[0021] Cluster the first embedding vectors based on a preset clustering algorithm according to the second parameter to obtain at least one vector cluster.
[0022] In some embodiments, the using the large language model to determine, from each second text, a text that matches the first theme view as the third text representing the event to be processed includes:
[0023] For each second text, input the second text and the first prompt word into the large language model to obtain the second theme view of each second text;
[0024] Calculate the similarity between the first theme view and the second theme view of each second text;
[0025] If the calculated similarity is greater than the preset similarity threshold, then determine the second text as the third text representing the event to be processed.
[0026] In some embodiments, the determining, by using the large language model, a text that matches the first theme view from each second text as the third text representing the event to be processed includes:
[0027] For each second text, input the second text, the first theme view, and the fourth prompt word into the large language model to obtain the matching degree of the second text; wherein, the fourth prompt word is used to indicate the matching degree between the text generated by the large language model and the theme view; if the matching degree of the second text is greater than the preset matching degree threshold, then determine the second text as the third text representing the event to be processed;
[0028] and / or
[0029] The obtaining of multiple texts containing event information of events as the first text includes:
[0030] Obtain multiple texts containing event information of events;
[0031] Preprocess each obtained text respectively to obtain multiple first texts; wherein, the preprocessing includes at least one of the following: denoising, stop word filtering, and stemming;
[0032] and / or
[0033] The method further includes:
[0034] Based on the occurrence time of the event characterized by the third text, use the time series analysis algorithm to obtain the analysis result of the occurrence period of the events in the category to be processed.
[0035] In some embodiments, after the determining, by using the large language model, a text that matches the first theme view from each second text as the third text representing the event to be processed, the method further includes:
[0036] Input the third text and the second prompt word into the large language model to obtain the title of the text in the category to be processed; wherein, the second prompt word is used to indicate the large language model to generate the title of the input text;
[0037] and / or
[0038] Input the third text and the fifth prompt into the large language model to obtain the cause of the event in the category to be processed; wherein, the fifth prompt is used to instruct the large language model to generate the cause of the event represented by the input text.
[0039] And / or,
[0040] Input the third text and the sixth prompt into the large language model to obtain the risk index of the event in the category to be processed; wherein, the sixth prompt is used to instruct the large language model to generate the risk index of the event represented by the input text, and the risk index of an event includes at least one of the following: the amount involved in the event, the number of people involved in the event, and the scope of the area affected by the event.
[0041] And / or,
[0042] Input the third text and the seventh prompt into the large language model to obtain the risk level of the event in the category to be processed; wherein, the seventh prompt is used to instruct the large language model to generate the risk level of the event represented by the input text, and the risk level of an event is determined based on the risk index of the event.
[0043] In the second aspect of the embodiments of the present application, an event recognition device is provided, and the device includes:
[0044] A first text acquisition module, configured to acquire multiple texts containing event information of events as the first text.
[0045] A second text determination module, configured to determine, based on a preset clustering algorithm, the text whose represented event belongs to the category to be processed from each first text as the second text.
[0046] A theme view generation module, configured to input each second text and the first prompt into the large language model to obtain the first theme view of the text in the category to be processed; wherein, the first prompt is used to instruct the large language model to generate the theme view of the input text.
[0047] A matching module, configured to use the large language model to determine, from each second text, the text that matches the first theme view as the third text representing the event to be processed.
[0048] In some embodiments, the second text determination module includes:
[0049] An encoding sub-module, configured to encode each first text to obtain a first embedding vector of each first text.
[0050] A clustering sub-module, configured to cluster the first embedding vectors based on a preset clustering algorithm to obtain at least one vector cluster.
[0051] A vector cluster selection sub-module, configured to select one vector cluster belonging to the category to be processed from the obtained at least one vector cluster;
[0052] A second text determination sub-module, configured to use the first text represented by the first embedding vector included in the selected vector cluster as the second text.
[0053] In some embodiments, the apparatus further includes:
[0054] A parameter acquisition module, configured to input the average length, standard deviation of each first text, and a third prompt word to a large language model before encoding each first text to obtain the first parameter for dimensionality reduction processing and the second parameter for clustering;
[0055] Wherein, the first parameter includes at least one of the following: the dimension of the embedded vector after dimensionality reduction, the minimum distance between each embedded vector after dimensionality reduction, and the number of adjacent points when constructing a neighborhood after dimensionality reduction; the second parameter includes at least one of the following: the minimum number of vectors included in a vector cluster, the minimum number of vectors in the neighborhood of the center of a vector cluster; the third prompt word is used to: instruct the large language model to generate parameters for dimensionality reduction processing and parameters for clustering based on the average length and standard deviation of the input text;
[0056] The encoding sub-module is specifically configured to:
[0057] Encode each first text and perform dimensionality reduction processing on the encoding result according to the first parameter to obtain the first embedding vector of each first text;
[0058] The clustering sub-module is specifically configured to:
[0059] Cluster the first embedding vectors according to the second parameter based on a preset clustering algorithm to obtain at least one vector cluster.
[0060] In some embodiments, the matching module is specifically configured to:
[0061] For each second text, input the second text and the first prompt word to the large language model to obtain the second topic view of each second text;
[0062] Calculate the similarity between the first topic view and the second topic view of each second text;
[0063] If the calculated similarity is greater than a preset similarity threshold, determine the second text as the third text representing the event to be processed.
[0064] In some embodiments, the matching module is specifically configured to:
[0065] For each second text, input the second text, the first theme view, and the fourth prompt word into the large language model to obtain the matching degree of the second text; wherein, the fourth prompt word is used to indicate the matching degree between the text generated by the large language model and the theme view;
[0066] If the matching degree of the second text is greater than the preset matching degree threshold, determine the second text as the third text representing the event to be processed;
[0067] And / or,
[0068] The first text acquisition module is specifically configured to:
[0069] Acquire multiple texts containing event information of events;
[0070] Preprocess each of the acquired texts to obtain multiple first texts; wherein, the preprocessing includes at least one of the following: denoising, stop word filtering, and stemming;
[0071] And / or,
[0072] The device further includes:
[0073] A period analysis module, configured to obtain an analysis result of the occurrence period of the events in the to-be-processed category by using a time series analysis algorithm based on the occurrence time of the events represented by the third text.
[0074] In some embodiments, the device further includes:
[0075] A title generation module, configured to input the third text and the second prompt word into the large language model to obtain a title of the text of the to-be-processed category; wherein, the second prompt word is used to indicate the large language model to generate a title of the input text;
[0076] And / or,
[0077] A cause of occurrence acquisition module, configured to, after determining, by using the large language model, a text that matches the first theme view from each second text as the third text representing the event to be processed, input the third text and the fifth prompt word into the large language model to obtain the cause of occurrence of the event in the to-be-processed category; wherein, the fifth prompt word is used to indicate the large language model to generate the cause of occurrence of the event represented by the input text;
[0078] And / or,
[0079] A risk index acquisition module, configured to, after using the large language model to determine, from each second text, a text that matches the first theme view as a third text representing a to-be-processed event, input the third text and a sixth prompt word into the large language model to obtain a risk index of the event in the to-be-processed category; wherein, the sixth prompt word is used to instruct the large language model to generate a risk index of the event represented by the input text, and a risk index of an event includes at least one of the following: the amount involved in the event, the number of people involved in the event, and the regional scope affected by the event;
[0080] and / or,
[0081] A risk level acquisition module, configured to, after using the large language model to determine, from each second text, a text that matches the first theme view as a third text representing a to-be-processed event, input the third text and a seventh prompt word into the large language model to obtain a risk level of the event in the to-be-processed category; wherein, the seventh prompt word is used to instruct the large language model to generate a risk level of the event represented by the input text, and a risk level of an event is determined based on the risk index of the event.
[0082] In a third aspect of the embodiments of the present application, an electronic device is provided, including:
[0083] A memory for storing a computer program;
[0084] A processor, configured to implement the event recognition method described in any one of the above when executing the program stored on the memory.
[0085] In yet another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, and when the computer program is executed by a processor, the event recognition method described in any one of the above is implemented.
[0086] In yet another aspect of the embodiments of the present application, a computer program product including instructions is provided, and when it runs on a computer, it causes the computer to execute the event recognition method described in any one of the above.
[0087] Advantageous effects of the embodiments of the present invention:
[0088] Based on the event recognition method provided in the embodiments of the present application, the first text may contain texts representing different events, and the categories to which different events belong may be different. Correspondingly, the electronic device can determine multiple texts (i.e., the second texts) representing events belonging to the category to be processed from the first text, and input each second text and a prompt word (i.e., the first prompt word) for indicating the topic view of the text for generating the input to the large language model. Correspondingly, the large language model can output the topic view of the text of the category to be processed (i.e., the first topic view). The first topic view of the text of the category to be processed can provide a concise description of the events represented by the text of the category to be processed. Therefore, after obtaining the first topic view of the text of the category to be processed, the electronic device can also use the large language model to determine the third text that matches the first topic view from each second text. That is to say, the third text with a higher occurrence frequency of the events represented in the first text can be determined. Correspondingly, the event represented by the third text is the event to be processed. In this way, the efficiency of identifying the event to be processed can be improved.
[0089] Of course, it is not necessary for any product or method implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0090] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.
[0091] Figure 1 The first flowchart of the event recognition method provided in the embodiments of the present application;
[0092] Figure 2 The second flowchart of the event recognition method provided in the embodiments of the present application;
[0093] Figure 3 The third flowchart of the event recognition method provided in the embodiments of the present application;
[0094] Figure 4 The schematic diagram of the clustering result obtained by clustering the first embedding vector of the first text provided in the embodiments of the present application;
[0095] Figure 5 The fourth flowchart of the event recognition method provided in the embodiments of the present application;
[0096] Figure 6 The structural diagram of an event recognition device provided in the embodiments of the present application;
[0097] Figure 7 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific embodiments
[0098] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art based on this application belong to the scope of protection of the present invention.
[0099] With the development of network technology, the public can upload opinions and views on a specified event on relevant platforms through the network.
[0100] For example, relevant departments (such as the traffic management department of a city) can establish a problem feedback platform for the public. Correspondingly, the public can upload the information of the problems to be reflected to this problem feedback platform. The information of the problems to be reflected can be an event that occurs in a certain area.
[0101] For another example, a certain online media can establish a feedback platform provided to netizens. Correspondingly, netizens can upload the information of the events that they hope the media will intervene in to this feedback platform.
[0102] However, the information received by the platform may be massive, and there may also be a large number of information on repeated events in the received information. Each piece of information may also contain invalid content (such as garbled characters) that has nothing to do with the event that the user needs to reflect. In one way, the received texts can be statistically analyzed one by one manually, and the texts representing the events can be determined. Subsequently, the users of the platform (such as the above-mentioned relevant departments or online media, etc.) can understand the event information of the event to be processed according to the texts representing the event to be processed, and process the event to be processed.
[0103] However, the efficiency of identifying events manually is relatively low.
[0104] An embodiment of the present application provides an event recognition method, and this method can be applied to an electronic device. For example, the electronic device can be the server of the above platform.
[0105] See Figure 1 , Figure 1 This is the first flowchart of the event recognition method provided by an embodiment of the present application, and this method includes the following steps:
[0106] S101: Obtain multiple texts containing event information of events as the first texts.
[0107] S102: Determine the texts whose represented events belong to the category to be processed from each first text based on a preset clustering algorithm as the second texts.
[0108] S103: Input each second text and the first prompt word into the large language model to obtain the first topic view of the texts of the category to be processed.
[0109] Among them, the first prompt word is used to instruct the large language model to generate the topic view of the input text.
[0110] S104: Use the large language model to determine the texts that match the first topic view from each second text as the third texts representing the events to be processed.
[0111] Based on the event recognition method provided in the embodiments of the present application, the first text may contain texts representing different events, and the categories to which different events belong may be different. Correspondingly, the electronic device can determine multiple texts (i.e., the second texts) whose represented events belong to the category to be processed from the first text, and input each second text and the prompt word (i.e., the first prompt word) for instructing the large language model to generate the topic view of the input text into the large language model. Correspondingly, the large language model can output the first topic view of the texts of the category to be processed (i.e., the first topic view). The first topic view of the texts of the category to be processed can provide a general and brief description of the events represented by the texts of the category to be processed. Therefore, after obtaining the first topic view of the texts of the category to be processed, the electronic device can also use the large language model to determine the third texts that match the first topic view from each second text. That is to say, based on a preset clustering algorithm, the third texts with a higher occurrence frequency of the represented events can be determined from the first text. Correspondingly, the events represented by the third texts are the events to be processed. In this way, the efficiency of recognizing the events to be processed can be improved.
[0112] Regarding step S101, the electronic device can obtain multiple texts of users (such as the above-mentioned community owners, Internet users, etc.). Among them, the content of each text is used to describe an event, that is, the text contains the event information of the event.
[0113] Among them, the texts containing the event information of the event obtained by the electronic device can be uploaded by users (such as the above-mentioned community owners, Internet users, etc.) through the above platform, or collected in advance through grid inspections.
[0114] For example, an electronic device can obtain 100,000 pieces of original text uploaded by a user from January 2022 to December 2022. Taking 23 pieces of the original text as an example, for instance, the original text 1 is: "The traffic signal setting at the intersection of Bing Avenue and Ding Road is unreasonable." The original text 2 is: "The elevator in Jia Community has a malfunction." The original text 3 is: "The implementation of garbage classification in Jia Community is ineffective." The original text 4 is: "The water supply in Yi Community is unstable." The original text 5 is: "The public transportation vehicle on Route N1 is late." The original text 6 is: "The construction site near Jia Community has noise during midnight construction." The original text 7 is: "There is a shortage of parking spaces in Jia Shopping Mall." The original text 8 is: Text content 8. The original text 9 is: Text content 9. The original text 10 is: Text content 10. The original text 11 is: Text content 11. The original text 12 is: Text content 12. The original text 13 is: Text content 13. The original text 14 is: Text content 14. The original text 15 is: Text content 15. The original text 16 is: Text content 16. The original text 17 is: Text content 17. The original text 18 is: Text content 18. The original text 19 is: Text content 19. The original text 20 is: Text content 20. The original text 21 is: Text content 21. The original text 22 is: Text content 22. The original text 23 is: Text content 23. Among them, one text content represents the event information of an event.
[0115] In one implementation manner, the electronic device can directly use the multiple obtained original texts as the first text.
[0116] In some embodiments, the electronic device can also preprocess the multiple obtained original texts respectively to improve the quality of the obtained first text. Refer to Figure 2 , Figure 2 which is the second flowchart of the event recognition method provided by the embodiments of the present application. On the basis of Figure 1 , step S101 includes:
[0117] S1011: Obtain multiple texts containing event information of events.
[0118] S1012: Preprocess each of the obtained texts respectively to obtain multiple first texts.
[0119] Among them, the preprocessing includes at least one of the following: denoising, stop word filtering, and stemming.
[0120] In the embodiments of the present application, after the electronic device obtains multiple texts containing event information of events (i.e., original texts), it can preprocess each of the obtained original texts respectively, and use the preprocessing results of the original texts as the first text. Among them, the preprocessing includes at least one of the following: denoising, stop word filtering, and stemming.
[0121] In one implementation, technicians can preset various preprocessing operations required for each original text, as well as the execution order between various preprocessing operations. In this application, the execution order between various preprocessing operations is not limited.
[0122] For example, for each original text, the electronic device can first perform denoising processing on the original text, then perform stop word filtering processing on the obtained denoising result, and finally perform stemming processing on the obtained stop word filtering result to obtain the preprocessing result of the original text. For another example, for each original text, the electronic device can also first perform stop word filtering processing on the original text, then perform denoising processing on the obtained stop word filtering result, and finally perform stemming processing on the obtained denoising result to obtain the preprocessing result of the original text.
[0123] Correspondingly, the electronic device can perform denoising processing on the text to be denoised. That is, delete the noise words in the text. Among them, the noise words include: special characters, numbers, punctuation marks, etc. Among them, the text to be denoised can be the original text, or the text to be denoised can also be obtained by performing other preprocessing operations on the original text except for denoising processing.
[0124] In this way, the electronic device can clean the text, that is, delete the noise in the text. Since the noise in the text (i.e., special characters, numbers, punctuation marks, etc.) is usually irrelevant to the event represented by the text. Therefore, it is possible to reduce the proportion of invalid information in the obtained first text, and thus improve the quality of the first text. Furthermore, it is possible to improve the effect of subsequent classification.
[0125] The electronic device can perform stop word filtering processing on the text to be filtered. That is, delete the stop words in the text. Among them, the stop words can be words without actual meaning, and / or words that may affect the subsequent classification process. That is to say, the stop words are: words irrelevant to the event information of the event represented by the text. Correspondingly, technicians can preset stop words according to actual needs. For example, the stop words can be common articles, prepositions, conjunctions, etc., such as: "of", "is", "in", etc. Among them, the text to be filtered can be the original text, or the text to be filtered can also be obtained by performing other preprocessing operations on the original text except for stop word filtering processing.
[0126] In this way, the electronic device can remove the stop words in the text (i.e., common articles, prepositions, conjunctions, etc.), and remove the words that appear frequently in the text but are irrelevant to the event information of the event represented by the text. It is also possible to reduce the proportion of invalid information in the obtained first text, and thus improve the quality of the first text. Furthermore, reduce the computational complexity of subsequent clustering, and improve the efficiency and effect of subsequent event recognition to be processed.
[0127] An electronic device can perform stemming processing on text to be standardized. That is, the text is tokenized, and the continuous original text is segmented into independent words or phrases to obtain the tokenization result of the text. Furthermore, words with a preset part of speech can be extracted from the tokenization result of the text and represented in a unified text representation form as the stemming result of the text (which can also be called the standardization result). For example, the electronic device can use a preset tokenization tool, such as the jieba tokenization tool, to perform tokenization processing on the text to obtain the tokenization result of the text. The preset part of speech can be nouns and verbs.
[0128] In this way, the electronic device can segment the text into the smallest semantic units (words). Since nouns and verbs in the text usually have practical meanings, performing stemming processing on the original text, obtaining nouns and verbs in the text and representing them in a unified text representation form can further improve the quality of the first text and facilitate subsequent feature extraction and analysis.
[0129] For example, taking the order of denoising, stop word filtering, and stemming as an example, the process of preprocessing the above original text 1 "The traffic signal setting at the intersection of Bing Avenue and Ding Road is unreasonable" is as follows: First, perform denoising processing on the original text 1, and the denoising result is "The traffic signal setting at the intersection of Bing Avenue and Ding Road is unreasonable". Second, perform stop word filtering processing on the obtained denoising result, and the stop word filtering result is "Bing Avenue Ding Road intersection traffic signal setting unreasonable". Finally, perform stemming processing on the obtained stop word filtering result, and the preprocessing result of the original text is: "Bing Avenue, Ding Road, intersection, traffic signal, setting, unreasonable". Among them, the "," in a preprocessing result is a tokenization symbol used to distinguish each word in the text, and the electronic device does not actually include this tokenization symbol in the preprocessing results of each text.
[0130] Correspondingly, the electronic device can obtain the preprocessing results of each original text as the first text.
[0131] Based on the above processing, after obtaining each original text, the electronic device can perform preprocessing operations such as denoising, stop word filtering, and stemming on each original text and represent them in a unified text representation form. In this way, the proportion of invalid information in the obtained first text can be reduced, the quality of the first text can be improved, the computational complexity of subsequent clustering can be reduced, and the efficiency and effect of subsequent event recognition to be processed can be improved.
[0132] For step S102, the electronic device may cluster each first text according to the category to which the event represented by each first text characterization belongs, determine the first text whose represented event belongs to the category to be processed as the second text. Accordingly, the events represented by the other first texts except the second text in the first texts do not belong to the category to be processed. The preset clustering algorithm may be the HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) algorithm, or the preset clustering algorithm may be the K-Means clustering algorithm.
[0133] Among them, the category to be processed is selected according to actual needs from the categories to which the events represented by each original text belong.
[0134] In some embodiments, based on the preset topic modeling technique, each first text may be classified by using text features (i.e., the embedding vectors of the text) to determine the category to which the event represented by each first text belongs. The preset topic modeling technique may be: BERTopic, a topic modeling technique based on the pre-trained model BERT (Bidirectional Encoder Representations from Transformers).
[0135] See Figure 3 , Figure 3 which is the third flowchart of the event recognition method provided by the embodiments of the present application. On the basis of Figure 1 , step S102 includes:
[0136] S1021: Encode each first text to obtain the first embedding vector of each first text.
[0137] S1022: Cluster the first embedding vectors based on the preset clustering algorithm to obtain at least one vector cluster.
[0138] S1023: Select a vector cluster belonging to the category to be processed from the at least one obtained vector cluster.
[0139] S1024: Use the first text represented by the first embedding vector included in the selected vector cluster as the second text.
[0140] In the embodiments of the present application, for each first text, the electronic device may use a pre-trained language model to encode the first text to obtain an encoding result of the first text. Among them, the pre-trained language model may be BERT (Bidirectional Encoder Representations from Transformers, bidirectional encoder representation based on Transformer), or the pre-trained language model may be TextCNN (Text Convolutional Neural Networks, text convolutional neural network).
[0141] It can be understood that for a first text, the first text is composed of multiple words. Therefore, the electronic device may first perform word embedding processing, that is, convert the words in the first text into word vectors. Furthermore, based on the respective word vectors corresponding to the first text, a vector representing the first text is constructed. That is, the electronic device may convert each first text into a form represented by a high-dimensional vector to obtain an embedding vector of each first text. The embedding vector of a first text can represent the semantic information of the first text.
[0142] For example, for the above-mentioned original text 1, the first text corresponding to the original text 1 (which can be called the first text 1) is: "Bing Avenue, Ding Road, intersection, traffic signal, setting, unreasonable". Furthermore, the electronic device may encode the first text 1 to obtain an encoding result of the first text 1, and the encoding result of the first text 1 is: [0.23, -0.11,..., 0.76].
[0143] Similarly, for the above-mentioned original text 2 "An elevator in Jia Community has a malfunction", the first text corresponding to the original text 2 (which can be called the first text 2) is: "Jia Community, elevator, malfunction". Furthermore, the electronic device may encode the first text 2 to obtain an encoding result of the first text 2, and the encoding result of the first text 2 is: [-0.34, 0.57,..., -0.09].
[0144] Similarly, for the above-mentioned original text 3 "The garbage classification in Jia Community is not effectively implemented", the first text corresponding to the original text 2 (which can be called the first text 3) is: "Jia Community, garbage, classification, implementation, ineffective". Furthermore, the electronic device may encode the first text 3 to obtain an encoding result of the first text 3, and the encoding result of the first text 3 is: [0.12, 0.36,..., -0.83]. And so on, the electronic device can obtain the encoding results of each first text, and the specific obtaining process will not be elaborated here.
[0145] In one implementation, after obtaining the encoding results of the first texts, the electronic device can directly determine the encoding results of the first texts as the first embedding vectors of the first texts. Correspondingly, the first embedding vector of a first text can represent the semantic information of the first text.
[0146] In another implementation, step S1021 described above includes:
[0147] Encode each of the first texts, and perform dimensionality reduction processing on the encoding results according to a first parameter to obtain the first embedding vectors of the first texts.
[0148] In the embodiments of the present application, after obtaining the encoding results of the first texts, for the encoding result of each first text, the electronic device can perform dimensionality reduction processing on the encoding result of the first text according to the first parameter based on a preset dimensionality reduction algorithm, embed the high-dimensional embedding vector (i.e., the encoding result of the first text) into a low-dimensional feature space, and thus can reduce the high-dimensional embedding vector to a lower dimension. Correspondingly, the electronic device can determine the dimensionality reduction result corresponding to the encoding result of the first text as the first embedding vector of the first text. Correspondingly, the first embedding vector of a first text can represent the semantic information of the first text. For example, the dimension of the first embedding vector can be two-dimensional or three-dimensional.
[0149] Among them, the preset dimensionality reduction algorithm can be: UMAP (Uniform Manifold Approximation and Projection) algorithm, or t-SNE (t-Distributed Stochastic Neighbor Embedding) algorithm.
[0150] The first parameter represents: the parameter used for dimensionality reduction processing. The first parameter includes at least one of the following: the dimension of the embedding vector after dimensionality reduction (which can be denoted as n_components), the minimum distance between the embedding vectors after dimensionality reduction (which can be denoted as min_dist), and the number of adjacent points when constructing a neighborhood after dimensionality reduction (which can be denoted as n_neighbors).
[0151] Among them, the dimension of the embedded vector after dimensionality reduction (n_components) represents the target dimension after dimensionality reduction of the encoding results of each first text. The minimum distance (min_dist) between the embedded vectors after dimensionality reduction is used to control the minimum distance of data points in the low-dimensional feature space after dimensionality reduction, which can reflect the compactness of the vector clusters. Here, one data point corresponds to one first embedded vector. The number of adjacent points (n_neighbors) when constructing the neighborhood after dimensionality reduction is used to control the degree of preservation of the global structure, and n_neighbors represents the number of neighbors considered for each data point when constructing the local neighborhood.
[0152] In one implementation, a technician can preset the parameter values of each first parameter. For example, n_neighbors can be set to 15, n_components can be set to 2, and min_dist can be set to 0.1.
[0153] In another implementation, the electronic device can use a large language model to obtain the parameter values of the first parameter. The specific process of obtaining the first parameter will be described in subsequent embodiments.
[0154] For example, for the encoding result of the above first text 1 (i.e., [0.23, -0.11,..., 0.76]), the electronic device can obtain the dimensionality reduction result corresponding to the encoding result of the first text 1 based on a preset dimensionality reduction algorithm as the first embedded vector of the first text 1. Similarly, for the encoding result of the above first text 2 (i.e., [-0.34, 0.57,..., -0.09]), the electronic device can obtain the dimensionality reduction result corresponding to the encoding result of the first text 2 based on a preset dimensionality reduction algorithm as the first embedded vector of the first text 2. And so on, the electronic device can obtain the first embedded vectors of each first text, and the specific obtaining process will not be elaborated here.
[0155] In this way, the electronic device can use a preset dimensionality reduction algorithm (such as UMAP) to perform dimensionality reduction processing on the encoding results of the first text. In this way, it is possible to reduce the dimension of the first embedded vectors of each first text, which can also reduce the amount of computation and memory usage required for subsequent operations, reduce the computational complexity, and facilitate visualization and subsequent clustering analysis. And in the subsequent clustering process, using data with a lower dimension for clustering can improve the efficiency and effect of clustering.
[0156] Regarding step S1022, after obtaining the first embedding vectors of the first texts, the electronic device may cluster the first embedding vectors based on a preset clustering algorithm to obtain at least one vector cluster. The preset clustering algorithm may be the HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) algorithm, or the preset clustering algorithm may be the K-Means clustering algorithm.
[0157] Correspondingly, since the first embedding vector of a first text can represent the semantic information of the first text. Therefore, the electronic device can use the preset clustering algorithm to group text vectors (i.e., first embedding vectors) with similar semantics into the same vector cluster based on the distribution among the first embedding vectors of the first texts. That is, in the feature space to which the first embedding vectors belong, the first embedding vectors in a region with a higher density represent the first embedding vectors in a vector cluster, and the first embedding vectors in a region with a lower density represent noise.
[0158] It can be understood that in the process of clustering the first embedding vectors based on the HDBSCAN algorithm, there is no need to pre-specify the number of vector clusters obtained. The HDBSCAN algorithm can automatically determine the number of vector clusters in the clustering result according to the density characteristics of the data (i.e., the first embedding vectors). In this way, the method for identifying a to-be-processed event provided in the embodiments of the present application is particularly applicable to application scenarios where the categories to which the events represented by the first texts need to be divided are unknown.
[0159] In addition, if there is a first text whose semantics are significantly different from those of other first texts, it can be marked as noise. That is, the HDBSCAN algorithm can effectively identify the noise in the data, which can ensure the accuracy of the clustering result.
[0160] In this way, in the process of obtaining the embedding vector of a first text, the electronic device can convert the discrete, high-dimensional, and sparse word vector representation into a continuous, low-dimensional, and dense vector representation. Correspondingly, the obtained first embedding vectors of the first texts can reflect the semantic relationship between words, which is convenient for subsequent clustering operations.
[0161] In some embodiments, step S1022 includes:
[0162] Cluster the first embedding vectors based on a preset clustering algorithm according to a second parameter to obtain at least one vector cluster.
[0163] In the embodiments of the present application, the second parameter represents: a parameter for clustering. The second parameter includes at least one of the following: the minimum number of vectors included in a vector cluster (which can be denoted as min_cluster_size), the minimum number of vectors within the neighborhood of the center of a vector cluster (which can be denoted as min_samples), and the minimum distance between clusters (which can be denoted as cluster_selection_epsilon).
[0164] Among them, the minimum number of vectors included in a vector cluster (min_cluster_size) represents: the minimum number of data points required to form a vector cluster. Among them, one data point corresponds to one first embedding vector. It can be understood that min_cluster_size represents the minimum scale of the vector cluster. If the value of min_cluster_size is relatively large, it may cause vector clusters with fewer data points to be regarded as noise.
[0165] The minimum number of vectors within the neighborhood of the center of a vector cluster (min_samples), that is, the minimum number of neighbors around the center (i.e., the core point) of each vector cluster. Min_samples is used to: control the identification of noise in data points. Correspondingly, if the value of min_samples is relatively large, it means that the density of data points required to form a vector cluster is relatively high, and it may cause the number of data points determined to be noise to increase.
[0166] The minimum distance between clusters (cluster_selection_epsilon) is used to control the merging of vector clusters. That is to say, if the distance between two vector clusters is less than cluster_selection_epsilon, these two vector clusters will be merged.
[0167] In one implementation, a technician can pre-set the parameter values of each second parameter. For example, min_cluster_size can be set to 5, min_samples can be set to none, and cluster_selection_epsilon can be set to 0.01.
[0168] In another implementation, the electronic device can use a large language model to obtain the parameter values of the second parameter. The specific process of obtaining the second parameter will be described in subsequent embodiments.
[0169] Based on the above processing, during the process of clustering the first embedding vectors of each first text, the electronic device can control the clustering result by setting parameters for clustering (i.e., the second parameters). That is to say, the electronic device can control the number and size of the vector clusters obtained based on the first embedding vectors of each first text according to actual needs. In this way, the flexibility of clustering can be ensured, and clustering can be performed according to actual needs.
[0170] For each vector cluster obtained according to the above steps S1021 to S1022, the electronic device can determine the first texts represented by the first embedding vectors included in the vector cluster as the first texts belonging to the same category.
[0171] That is to say, the electronic device can determine the classification result of the events represented by each first text according to the clustering result of the first embedding vectors of each first text. Correspondingly, the number and size of the vector clusters obtained during the clustering process are used to reflect the classification granularity for classifying the events represented by the first text. That is, technicians can set the classification granularity when classifying events by setting the parameters of clustering.
[0172] It can be understood that for the obtained first text, if the classification granularity of the event is finer, the number of vector clusters obtained by clustering is more, that is, the number of categories to which the events represented by each first text belong is more. Correspondingly, the number of first embedding vectors included in a vector cluster is less, and the number of first texts corresponding to the same category is less. If the classification granularity of the event is coarser, the number of vector clusters obtained by clustering is less, that is, the number of categories to which the events represented by each first text belong is less. Correspondingly, the number of first embedding vectors included in a vector cluster is more, and the number of first texts corresponding to the same category is more.
[0173] For example, taking the above original text 1, original text 5, original text 9, original text 18, and original text 19 as examples, original text 1 is "The traffic signal setting at the intersection of Bing Avenue and Ding Road is unreasonable"; original text 5 is "The N1 bus is late"; original text 9 is text content 9; original text 18 is text content 18; original text 19 is text content 19.
[0174] In the case where the classification granularity of events is relatively coarse, the first embedding vectors included in a vector cluster (which can be denoted as vector cluster A) include: the first embedding vector of the first text corresponding to the above-mentioned original text 1 (i.e., the first text 1), the first embedding vector of the first text corresponding to the original text 5 (i.e., the first text 5), the first embedding vector of the first text corresponding to the original text 9 (i.e., the first text 9), the first embedding vector of the first text corresponding to the original text 18 (i.e., the first text 18), and the first embedding vector of the first text corresponding to the original text 19 (i.e., the first text 19). At this time, the events represented by the above-mentioned first text 1, first text 5, first text 9, first text 18, and first text 19 belong to the same category (this category can be denoted as: traffic).
[0175] For another example, in the case where the classification granularity of events is relatively fine, the first embedding vectors included in a vector cluster (which can be denoted as vector cluster B) include: the first embedding vector of the first text 1, the first embedding vector of the first text 5, the first embedding vector of the first text 9, and the first embedding vector of the first text 18. At this time, the events represented by the above-mentioned first text 1, first text 5, first text 9, and first text 18 belong to the same category (this category can be denoted as: city - traffic).
[0176] Correspondingly, since one vector cluster corresponds to one category, the electronic device can select one vector cluster from the obtained at least one vector cluster. Correspondingly, the category corresponding to this vector cluster is the category to be processed. Furthermore, the first text represented by the first embedding vectors included in the selected vector cluster is used as the second text. That is, the second text is the first text whose represented event belongs to the category to be processed.
[0177] In one implementation manner, for each vector cluster, the electronic device can determine the top first number of words (i.e., keywords) with the highest occurrence frequency among the first texts corresponding to the respective first embedding vectors included in this vector cluster, and obtain the keyword set of this vector cluster. For example, the first number can be 2.
[0178] For example, when the first number is 2, for the above-mentioned vector cluster A, among the first texts corresponding to the respective first embedding vectors included in vector cluster A, the top 2 words (i.e., keywords) with the highest occurrence frequency are "road" and "congestion". Correspondingly, the keyword set of vector cluster A is: road - congestion.
[0179] In another implementation manner, before step S1023, the method further includes:
[0180] Step 1: For each vector cluster, based on a preset keyword extraction algorithm, extract multiple keywords from the first texts corresponding to the respective first embedding vectors included in this vector cluster, and use them as the keyword set of this vector cluster.
[0181] Correspondingly, step S1023 includes:
[0182] Based on the keyword sets of each vector cluster, select one vector cluster belonging to the category to be processed from the obtained at least one vector cluster.
[0183] In the embodiment of the present application, one vector cluster corresponds to one category. That is to say, the electronic device can determine the first text belonging to the same category according to one vector cluster. Correspondingly, for each obtained vector cluster, the electronic device can extract multiple keywords from each first text corresponding to each first embedding vector included in the vector cluster based on a preset keyword extraction algorithm, as the keyword set of the vector cluster.
[0184] Among them, the preset keyword extraction algorithm can be: C-TF-IDF (Class-based Term Frequency-Inverse Document Frequency) algorithm. The C-TF-IDF algorithm is extended on the basis of the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm, and the C-TF-IDF algorithm is used to extract representative keywords from the texts of each category.
[0185] That is to say, the C-TF-IDF algorithm can identify the words most representative of the text of this category, rather than just high-frequency words. In this way, the frequency of words in a specific category can be considered in the process of determining the keyword sets of each vector cluster, and the interference of general words on the recognition result can be reduced. If the TF-IDF algorithm is used for recognition, some words that frequently appear in multiple categories (i.e., general words) may be used as keywords. The C-TF-IDF algorithm suppresses the weight of general words through the inverse class frequency, improving the uniqueness and recognition of the keywords of each vector cluster obtained.
[0186] In one implementation, the above step 1 includes:
[0187] For each vector cluster, according to the third parameter, based on the preset keyword extraction algorithm, extract multiple keywords from each first text corresponding to each first embedding vector included in the vector cluster, as the keyword set of the vector cluster.
[0188] In the embodiment of the present application, each first text corresponding to each first embedding vector included in one vector cluster can be called a theme. Correspondingly, obtaining the keyword set of a vector cluster is also obtaining the keyword set of the theme corresponding to the vector cluster.
[0189] The third parameter represents: the parameter for keyword extraction. The third parameter includes at least one of the following: the phrase range for keyword extraction (which can be denoted as n_gram_range), the number of keywords in the keyword set for each topic (which can be denoted as top_n_words), the minimum number of texts included in each topic (which can be denoted as min_topic_size), the setting method for the number of topics (which can be denoted as nr_topics), and whether to use BM25 weighting for keyword extraction (which can be denoted as bm25_weighting).
[0190] Among them, the phrase range for keyword extraction (n_gram_range) is used to define the phrase range considered when extracting keywords. For example, if the value of the phrase range for keyword extraction is (1, 2), it means that the extracted keywords can be single words and phrases composed of two words. The number of keywords in the keyword set for each topic (top_n_words) represents the number of keywords with higher weights extracted for each topic. The higher the weight of a keyword, the more important the keyword is relative to the topic. The minimum number of texts included in each topic (min_topic_size) represents the minimum value among the numbers of the first texts included in each topic, that is, the minimum scale of the topic. When the setting method for the number of topics (nr_topics) is auto (automatic), it means that the number of topics is automatically optimized according to the data. When whether to use BM25 weighting for keyword extraction (bm25_weighting) is False (no), it means that BM25 weighting is not used. BM25 (Best Matching 25) is an information retrieval algorithm.
[0191] In one implementation, a technician can preset the parameter values of each third parameter. For example, n_gram_range can be set to (1, 2), top_n_words can be set to 10, min_topic_size can be set to 10, nr_topics can be set to auto, and bm25_weighting can be set to False.
[0192] In another implementation, the electronic device can use a large language model to obtain the parameter values of the third parameter. The specific process of obtaining the third parameter will be described in subsequent embodiments.
[0193] For example, for the above-mentioned vector cluster A, the electronic device can extract multiple keywords from each first text corresponding to each first embedding vector included in the vector cluster as the keyword set of the vector cluster based on a preset keyword extraction algorithm. For example, the keyword set of the obtained vector cluster A can be: road - congestion.
[0194] In this way, the electronic device can extract multiple keywords from each first text corresponding to each first embedding vector included in a vector cluster based on BERTopic, and use them as the keyword set of the vector cluster. In this way, the category corresponding to each vector cluster can be automatically determined.
[0195] Furthermore, after obtaining the keyword sets of each vector cluster, the managers of the relevant platform (such as the above-mentioned online media, community property) can select one from all the vector clusters through the keyword sets of each vector cluster. Correspondingly, the electronic device can use the first text represented by the first embedding vector included in the selected vector cluster as the second text.
[0196] Based on the above processing, the electronic device can process the first text to obtain the first embedding vector of each first text. It realizes the conversion of subtraction Chinese characters into vector form for representation. Furthermore, clustering processing can be performed on the first embedding vectors, and thus the first text can be classified through each vector cluster obtained by clustering, and the text whose represented event belongs to the category to be processed is determined as the second text.
[0197] See Figure 4 , Figure 4 which is a schematic diagram of the clustering result obtained by clustering the first embedding vectors of the first text provided by the embodiment of the present application. If the number of first texts is 21, that is, the number of first embedding vectors is 21. Figure 4 In [the figure], the area corresponding to each dotted circle frame corresponds to a vector cluster, and each black dot corresponds to the first embedding vector of a first text. The position of the black dot in the rectangular frame represents the position of the first embedding vector of the first text in the feature space. The clustering result obtained by clustering the first embedding vectors of the 21 first texts can be shown in Table 1.
[0198] Table 1
[0199]
[0200] Among them, one row of data in the table corresponds to a vector cluster, and the cluster number represents the number of the vector cluster. The text serial number represents the serial number of each first text corresponding to each first embedding vector in the vector cluster. The keyword represents the keyword of the vector cluster obtained by extraction.
[0201] The vector cluster represented by Region 1 has a cluster number of 1. This vector cluster contains the first embedding vectors of the first texts with text serial numbers 1, 5, and 6, and the keyword of this vector cluster obtained by extraction is traffic.
[0202] The vector cluster represented by Region 2 has a cluster number of 2. This vector cluster contains the first embedding vectors of the first texts with text serial numbers 2, 7, 8, and 9, and the keyword of this vector cluster obtained by extraction is environmental sanitation.
[0203] The cluster number of the vector cluster represented by Region 3 is 3. The vector cluster contains the first embedding vectors of the first text with text numbers 3, 10, 11, 14, 15, and 18. The keyword extracted from this vector cluster is public facilities.
[0204] The cluster number of the vector cluster represented by Region 4 is 4. The vector cluster contains the first embedding vectors of the first text with text numbers 4, 12, 13, 16, 17, 19, and 20. The keyword extracted from this vector cluster is food safety.
[0205] The cluster number of the vector cluster represented by Region -1 is -1. The vector cluster contains the first embedding vector of the first text with text number 21. Since the distance between this first embedding vector and other first embedding vectors is relatively far, this first embedding vector can be determined as noise. Correspondingly, the electronic device can set the keyword of this vector cluster as noise.
[0206] It can be understood that in the clustering result, the event represented by the first text corresponding to the first embedding vector determined as noise has a low occurrence frequency. Correspondingly, the event corresponding to this noise is not the event to be processed in this application.
[0207] Regarding step S103, after obtaining the second text of the category to be processed, the electronic device can use a large language model to generate the topic view (i.e., the first topic view) of the second text of the category to be processed. That is, the electronic device can input each second text and the first prompt word into the large language model (LLM, Large Language Model).
[0208] Among them, the large language model can be ChatGPT (Chat Generative Pre-trained Transformer, chat generation type pre-trained transformation model), QWEN (Tongyi) model, etc. The large language model can learn the grammar and semantic rules of natural language, and it can perform corresponding text generation functions according to the input prompt word. The prompt word (prompt) refers to the initial text or instruction input into the model. In this application, the prompt word is the injection instruction input into the large language model, and the representation form of the prompt word is not limited to words, and can also be sentences, pictures, website addresses, etc.
[0209] The first prompt word is used to instruct the large language model to generate the topic view of the input text. For example, the first prompt word can be: "Please output the topic view of these multiple texts."
[0210] In one implementation, during the process of using the large language model to generate the first topic view of the second text of the category to be processed, the first prompt word can also include: the content requirement for the first topic view.
[0211] For example, the content requirements for the first topic view may include at least one of the following: disambiguation and enhancing semantic associations, etc.
[0212] Among them, disambiguation means: using a large language model to check for possible ambiguities in the generated first topic view to ensure the clarity and accuracy of the generated first topic view.
[0213] Enhancing semantics includes: association expansion and context supplementation. Association expansion means: identifying the associations between the first topic view and concepts such as events and people to enrich the connotation of the first topic view. Context supplementation means: adding context information such as time, place, and reason to the generated first topic view to make the first topic view more complete.
[0214] The large language model can obtain the first topic view deepened at the semantic level through semantic analysis. Additionally, add the required information (such as time, place, etc.) to the generated first topic view according to actual needs to enhance the depth and breadth of the first topic view.
[0215] Correspondingly, the large language model can perform multi-angle analysis on the multiple second texts according to the actual needs represented by the first prompt word, generate and output the first topic view of the text of the category to be processed. Correspondingly, the electronic device can obtain the first topic view with enhanced semantics, the first topic view of the text of the category to be processed, to ensure the accuracy of subsequent analysis.
[0216] In this way, the large language model can conduct in-depth analysis on the text of the category to be processed and extract the first topic view that can represent the text of the category to be processed. That is to say, the large language model can automatically generate refined and accurate first topic views from a large amount of text. In this way, it can significantly improve the accuracy and richness of the description of the first topic view and provide more powerful support for subsequent processing.
[0217] Regarding step S104, texts of the same category represent the same event. The event represented by the second text is the to-be-processed event with a higher occurrence frequency obtained by clustering among all the events represented by the first texts. It can be understood that among the second texts of the category to be processed, there may be a small number of texts with a low correlation degree or poor quality with the first topic view. Therefore, the electronic device can use the first topic view to screen the second texts of the category to be processed to determine the texts that match the first topic view from each second text as the third text.
[0218] The electronic device can determine the third text that matches the first topic view from each second text through any one of the following method 1 and method 2.
[0219] Method 1:
[0220] In some embodiments, step S104 includes:
[0221] Step S1041: For each second text, input the second text and the first prompt into the large language model to obtain the second topic view of each second text.
[0222] Step S1042: Calculate the similarity between the first topic view and the second topic view of each second text.
[0223] Step S10, if the calculated similarity is greater than the preset similarity threshold, then determine the second text as the third text representing the event to be processed.
[0224] In the embodiments of the present application, for each second text, the electronic device can use the large language model to generate the topic view of the second text (which can be called the second topic view). That is, the electronic device can input each second text and the first prompt into the large language model. Correspondingly, the large language model can process the second text, generate and output the second topic view of the second text.
[0225] Furthermore, the electronic device can calculate the similarity between the first topic view and the second topic view of the second text.
[0226] For example, the electronic device can obtain the encoding result of the first topic view and the encoding result of the second topic view of the second text based on the above-mentioned BERT (Bidirectional Encoder Representations from Transformers, bidirectional encoder representation based on Transformer). Furthermore, according to the distance between the encoding result of the first topic view and the encoding result of the second topic view of the second text, calculate the similarity between the first topic view and the second topic view of the second text.
[0227] If the calculated similarity is greater than the preset similarity threshold, it indicates that the semantics represented by the second topic view of the second text is close to the semantics represented by the first topic view, which can also indicate that the semantics represented by the second text is close to the semantics represented by the first topic view, that is, the second text matches the first topic view. Therefore, the electronic device can determine the second text as the third text.
[0228] If the calculated similarity is not greater than the preset similarity threshold, it indicates that the semantics of the second topic view representation of the second text is not close to the semantics of the first topic view representation, which can also indicate that the semantics of the second text representation is not close to the semantics of the first topic view representation, that is, the second text does not match the first topic view. Therefore, the electronic device will not determine the second text as the third text.
[0229] Based on the above processing, the electronic device can use the large language model to generate the second topic view of each second text. Furthermore, by matching the second topic views of each second text with the first topic view corresponding to the category to be processed, it is determined whether each second text matches the first topic view. In this way, texts in the second text with low relevance or poor quality to the first topic view corresponding to the category to be processed can be removed. Thus, the quality of the third text used to generate the title can be improved, and the accuracy of the generated title can be improved.
[0230] Method 2:
[0231] In some embodiments, step S104 includes:
[0232] S104a: For each second text, input the second text, the first topic view, and the fourth prompt word into the large language model to obtain the matching degree of the second text.
[0233] Among them, the fourth prompt word is used to instruct the large language model to generate the matching degree between the text and the topic view.
[0234] S104b: If the matching degree of the second text is greater than the preset matching degree threshold, then determine the second text as the third text representing the event to be processed.
[0235] In the embodiment of the present application, for each second text, the electronic device can use the large language model to determine the matching degree between the second text and the first topic view. That is, the electronic device can input the second text, the first topic view, and the fourth prompt word into the large language model.
[0236] Among them, the fourth prompt word is used to instruct the large language model to generate the matching degree between the text and the first topic view. For example, the fourth prompt word can be: "Please output the matching degree between this text and this first topic view".
[0237] In one implementation, before generating the matching degree between each second text and the first topic view using the large language model, the electronic device can also input into the large language model: multiple first sample texts, sample topic views, and the labels of each first sample text, and the label of a first sample text represents the matching degree of the first sample text corresponding to the sample topic view.
[0238] In the embodiments of the present application, the multiple first sample texts include: first sample texts with high quality and high relevance to the sample theme view, and first sample texts with low quality and low relevance to the sample theme view. Correspondingly, the matching degree represented by the label of the first sample text with high quality and high relevance to the sample theme view is greater than the matching degree represented by the label of the first sample text with low quality and low relevance to the sample theme view.
[0239] Among them, the quality of a text is determined based on the clarity of the text and the amount of information in the text.
[0240] The clarity of a text indicates the clarity and accuracy of the event represented by the text. The amount of information in a text indicates how much valuable information is contained in the text.
[0241] The relevance of a text to the theme view indicates the matching degree between the semantics represented by the content of the text and the semantics of the theme view.
[0242] That is to say, the large language model can learn how to evaluate the matching degree between a text and a theme view from multiple dimensions (i.e., quality and relevance) based on the input multiple first sample texts, the sample theme view, and the labels of each first sample text.
[0243] Correspondingly, for each second text input to the large language model, the large language model can process the second text and the first theme view to generate the matching degree between the second text and the first theme view.
[0244] Furthermore, the electronic device can screen the second text based on a preset matching degree threshold. That is, if the matching degree of a second text is greater than the preset matching degree threshold, it indicates that the second text matches the first theme view, and the second text is suitable as the text for generating the title corresponding to the category to be processed. Therefore, the electronic device can determine the second text as the third text.
[0245] If the matching degree of a second text is not greater than the preset matching degree threshold, it indicates that the quality of the second text is poor or the relevance to the first theme view is low, and the second text is not suitable as the third text. For example, in the second text, the text with poor quality or low relevance to the first theme view may be: the text with semantics inconsistent with the semantics represented by the first theme view, or the text with ambiguity and misleading.
[0246] Based on the above processing, the electronic device can use the large language model to perform concept consistency verification, that is, use the large language model to verify each second text to determine whether the second text belongs to or effectively expresses the first theme view, and obtain the matching degree between the second text and the first theme view. Furthermore, the matching degrees of the second texts are used for screening to filter out texts with a relatively low correlation or poor quality with the first theme view, ensuring the quality of the obtained third text.
[0247] In addition, by evaluating and screening the second texts, the data quality can be improved, ensuring that the retained third texts are highly relevant and of excellent quality, providing a reliable basis for subsequent processing. At the same time, filtering out texts with a low correlation or poor quality with the first theme view strengthens the consistency of the texts in the third text and reduces the impact of noise on subsequent processing.
[0248] In some embodiments, after step S104, the method further includes:
[0249] Input the third text and the second prompt into the large language model to obtain the title of the text of the category to be processed.
[0250] Wherein, the second prompt is used to instruct the large language model to generate the title of the input text.
[0251] In the embodiments of the present application, the electronic device can generate a title based on the third text as the title of the text of the category to be processed. That is, the electronic device can input the third text and the second prompt into the large language model.
[0252] Wherein, the second prompt is used to instruct the large language model to generate the title of the input text. For example, the second prompt can be: "Please output the title of these multiple texts".
[0253] In one implementation, during the process of using the large language model to generate the title of the third text of the category to be processed, the second prompt can also include: requirements for the title. For example, the requirements for the title can include restrictions on the number of words in the title and content requirements represented by the title. For example, the content requirements represented by the title are used to limit: the title contains concepts such as the result of an event, a person, etc.
[0254] Correspondingly, the large language model can perform multi-angle analysis on the multiple third texts according to the actual requirements represented by the second prompt, generate and output the titles of the multiple third texts. Correspondingly, the electronic device can obtain the title as the title of the text of the category to be processed.
[0255] Based on the event recognition method provided in this application, the electronic device can classify the obtained text based on the BERTopic technology. Since the BERTopic technology can achieve model parameter self-adaptation, that is, automatically adjust the clustering parameters according to the data volume of the text, the clustering effect can be optimized, and the clustering accuracy and stability can be improved.
[0256] Furthermore, by using the large language model to process the texts of the same category (category to be processed) to generate the titles of the texts of the category to be processed, the powerful semantic understanding ability of the large language model can be utilized to conduct in-depth semantic analysis on the obtained texts, significantly improving the accuracy and efficiency of the recognition of the events to be processed.
[0257] Compared with the method of directly inputting all texts into the large language model and classifying and recognizing the events to be processed through the large language model, the event recognition method provided in this application can improve the efficiency of the recognition of the events to be processed and reduce the computational amount of the large language model.
[0258] In addition, the event recognition method provided in the application introduces a concept verification and data screening link, that is, using the large language model to determine the third texts that match the first theme view from each second text. In this way, high-quality and highly relevant data can be screened out, noise and irrelevant data can be removed, ensuring the quality and consistency of the texts used to generate the titles, and thus ensuring the accuracy of the obtained titles.
[0259] In an embodiment, refer to Figure 5 , Figure 5 which is the fourth flowchart of the event recognition method provided in the embodiment of this application. On the basis of Figure 3 , before step S1021, the method further includes:
[0260] S1025: Input the average length, standard deviation of each first text, and the third prompt word into the large language model to obtain the first parameter for dimensionality reduction processing and the second parameter for clustering.
[0261] Among them, the first parameter includes at least one of the following: the dimension of the embedded vector after dimensionality reduction, the minimum distance between the embedded vectors after dimensionality reduction, and the number of adjacent points when constructing the neighborhood after dimensionality reduction; the second parameter includes at least one of the following: the minimum number of vectors included in a vector cluster, the minimum number of vectors in the neighborhood of the center of a vector cluster; the third prompt word is used to: instruct the large language model to generate the parameters for dimensionality reduction processing and the parameters for clustering based on the average length and standard deviation of the input text.
[0262] Step S1021 includes:
[0263] S1021a: Encode each first text and perform dimensionality reduction on the encoding result according to the first parameter to obtain the first embedding vector of each first text.
[0264] Step S1022 includes:
[0265] S1022a: Cluster the first embedding vectors based on a preset clustering algorithm according to the second parameter to obtain at least one vector cluster.
[0266] In the embodiments of the present application, the electronic device can utilize a large language model to obtain the first parameter for dimensionality reduction processing and the second parameter for clustering.
[0267] Correspondingly, the electronic device can calculate the average length and standard deviation of each first text. Furthermore, the average length, standard deviation of each first text, and the third prompt word are input into the large language model.
[0268] Among them, the third prompt word is used to: instruct the large language model to generate the parameter for dimensionality reduction processing and the parameter for clustering based on the average length and standard deviation of the input text.
[0269] In one implementation, before using the large language model to generate the parameter for dimensionality reduction processing and the parameter for clustering, the electronic device can also input into the large language model: the average length and standard deviation of multiple groups of second sample texts, and the labels corresponding to each group of second sample texts; where one group of second sample texts includes multiple sample texts, and the label corresponding to one group of second sample texts includes: the parameter values of each first parameter when performing dimensionality reduction on the embedding vectors of this group of second sample texts, and the parameter values of each second parameter when performing clustering on the embedding vectors of this group of second sample texts.
[0270] In addition, the label corresponding to one group of second sample texts can also include: the parameter values of each third parameter when extracting keywords from the embedding vectors of this group of second sample texts.
[0271] For example, the average length of a group of second sample texts is large, and the standard deviation is also large. This indicates that: the amount of data in the texts of this group of second sample texts is large and the event information represented is more complex. That is, the number of texts included in this group of second sample texts is large, and the categories to which they belong are diverse and subdivided. Correspondingly, among the first parameters of this group of second sample texts, the minimum distance (min_dist) between the embedded vectors after dimensionality reduction can be 0.1 - 0.3 to ensure a certain separation between the obtained vector clusters; the number of adjacent points (n_neighbors) when constructing neighborhoods after dimensionality reduction can be increased, such as it can be 30 - 50, to capture more global structures. Among the second parameters of this group of second sample texts, the minimum number of vectors included in a vector cluster (min_cluster_size) can be increased, such as it can be 10 - 20, to avoid treating smaller vector clusters as noise. The minimum number of vectors in the neighborhood of the center of a vector cluster can be increased, which can be adjusted according to the data density to reduce the number of embedded vectors determined as noise. Among the third parameters of this group of second sample texts, the minimum number of texts included in each topic (min_topic_size) can be increased, such as it can be 20 - 50, to ensure that the topics have sufficient text support. The number of topics can be set to auto to automatically optimize the number of topics based on the algorithm. The number of keywords (top_n_words) in the keyword set of each topic can be increased, such as it can be 15 - 20, to ensure that the obtained keywords are more comprehensive.
[0272] For example, the average length of a group of second sample texts is small, and the standard deviation is also small. This indicates that the amount of data in the texts of this group of second sample texts is small and the event information represented is relatively simple. That is, the number of texts included in this group of second sample texts is small, and the categories to which they belong are few and obvious. Correspondingly, among the first parameters of this group of second sample texts, the minimum distance (min_dist) between the embedded vectors after dimensionality reduction can remain unchanged or decrease, such as it can be 0.0 - 0.1, to ensure that the obtained vector clusters are closer; the number of adjacent points (n_neighbors) when constructing the neighborhood after dimensionality reduction can be reduced, such as it can be 5 - 15, to emphasize the local structure. Among the second parameters of this group of second sample texts, the minimum number of vectors (min_cluster_size) included in a vector cluster can be reduced, such as it can be 3 - 5, to capture and obtain more categories with a smaller scope. The minimum number of vectors in the neighborhood of the center of a vector cluster can be reduced to allow more vectors to be clustered into one vector cluster. Among the third parameters of this group of second sample texts, the minimum number of texts (min_topic_size) included in each topic can be reduced, such as it can be 5 - 10, to allow the texts to be subdivided into multiple categories. The number of topics can be set to auto, or alternatively, a smaller number of categories can be set according to actual needs. The number of keywords (top_n_words) in the keyword set of each topic can remain unchanged or decrease, such as it can be 10 - 15.
[0273] Correspondingly, after obtaining the first parameters for dimensionality reduction processing and the second parameters for clustering, the electronic device can encode each first text and perform dimensionality reduction processing on the encoding result according to the first parameters to obtain the first embedded vectors of each first text. Furthermore, according to the second parameters, based on a preset clustering algorithm, the first embedded vectors are clustered to obtain at least one vector cluster. The processes of dimensionality reduction and clustering can refer to the relevant descriptions in steps S1022 and S1023 above and will not be elaborated here.
[0274] Based on the above processing, the large language model can learn to determine the first parameters for dimensionality reduction processing, the second parameters for clustering, and the third parameters for keyword extraction based on the average length and standard deviation of the texts. Correspondingly, the electronic device can use the large language model to obtain the first parameters, second parameters, and third parameters suitable for processing the embedded vectors of the first text. In this way, the first text can be classified according to actual needs, and further, the accuracy and consistency of the determined second text can be improved.
[0275] In some embodiments, after step S104, the method further includes:
[0276] Input the third text and the fifth prompt into the large language model to obtain the cause of the event in the category to be processed.
[0277] Among them, the fifth prompt is used to instruct the large language model to generate the cause of the event represented by the input text.
[0278] In the embodiments of the present application, since the events represented by each third text belong to the category to be processed, the electronic device can use the large language model to process the third text to obtain the cause of the event in the category to be processed.
[0279] That is, the electronic device can input the third text and the fifth prompt into the large language model.
[0280] Among them, the fifth prompt is used to instruct the large language model to generate the cause of the event represented by the input text. For example, the fifth prompt can be: "Please output the cause of the event in this category."
[0281] Correspondingly, the large language model can process and analyze each third text, obtain and output the cause of the event represented by each third text. Furthermore, the electronic device can obtain the cause of the event in the category to be processed.
[0282] Based on the above processing, for the events in the category to be processed, the electronic device can use the large language model to analyze the third text corresponding to the category to be processed, clarify the cause of the events in the category to be processed, and obtain the core contradictions and demands of the events in the category to be processed.
[0283] In some embodiments, after step S104, the method further includes:
[0284] Input the third text and the sixth prompt into the large language model to obtain the risk index of the event in the category to be processed.
[0285] Among them, the sixth prompt is used to instruct the large language model to generate the risk index of the event represented by the input text. The risk index of an event includes at least one of the following: the amount involved in the event, the number of people involved in the event, and the geographical scope affected by the event.
[0286] In the embodiments of the present application, since the events represented by each third text belong to the category to be processed, the electronic device can use the large language model to process the third text to obtain the risk index of the event in the category to be processed. Among them, the risk index of an event includes at least one of the following: the amount involved in the event, the number of people involved in the event, and the geographical scope affected by the event.
[0287] That is, the electronic device can input the third text and the sixth prompt into the large language model.
[0288] Among them, the sixth prompt word is used to indicate the risk index of the event represented by the input text for the large language model to generate. For example, the sixth prompt word can be: "Please output the risk index of the event of this category."
[0289] Correspondingly, the large language model can process and analyze each third text, obtain and output the risk index of the event represented by each third text. Furthermore, the electronic device can obtain the risk index of the event in the category to be processed.
[0290] It can be understood that for events of different categories, the risk index of the event may be different. For example, for events of the natural disaster category, the risk index of the event can be the number of people involved in the event, the area affected by the event, etc.
[0291] Based on the above processing, for the event in the category to be processed, the electronic device can use the large language model to analyze the third text corresponding to the category to be processed, clarify the risk index of the event in the category to be processed, and facilitate subsequent analysis and processing.
[0292] In some embodiments, after step S104, the method further includes:
[0293] Input the third text and the seventh prompt word into the large language model to obtain the risk level of the event in the category to be processed.
[0294] Among them, the seventh prompt word is used to indicate the risk level of the event represented by the input text for the large language model to generate, and the risk level of an event is determined based on the risk index of the event.
[0295] In the embodiments of the present application, since the events represented by each third text belong to the category to be processed, the electronic device can use the large language model to process the third text to obtain the risk level of the event in the category to be processed. The risk level of an event is determined based on the risk index of the event.
[0296] That is, the electronic device can input the third text and the seventh prompt word into the large language model.
[0297] Among them, the seventh prompt word is used to indicate the risk level of the event represented by the input text for the large language model to generate. For example, the seventh prompt word can be: "Please determine the risk level of the event of this category based on the risk index of the event of this category."
[0298] Correspondingly, the large language model can process and analyze each third text, obtain and output the risk level of the event represented by each third text.
[0299] In one implementation, the large language model can combine real-time data and knowledge in the field to which the type of event belongs to dynamically analyze the identified risk indicators, evaluate the urgency and immediacy of the event, and determine the risk level of the event of this category. For example, for events with relatively high urgency and immediacy, the risk level of the event of this category can be determined to be high; for events with relatively low urgency and immediacy, the risk level of the event of this category can be determined to be low.
[0300] Based on the above processing, for different types of events, the large language model can identify specific risk indicators, and combine real-time data and domain knowledge for dynamic risk analysis to accurately evaluate the risk level of the event. This provides support for the subsequent rapid response and resource scheduling for events with a high risk level, and reduces the impact caused by high-risk events.
[0301] In this way, high-risk events can be quickly determined. Subsequently, the managers of the relevant platform can give priority to paying attention to and handling events of this category, realizing continuous monitoring and analysis of the data in the relevant platform, and achieving early warning and rapid response for high-risk events.
[0302] In some embodiments, the method further includes:
[0303] Based on the occurrence time of the event characterized by the third text, using a time series analysis algorithm, an analysis result of the occurrence period of the events in the category to be processed is obtained.
[0304] In the application embodiment, the electronic device can obtain the occurrence time of the event characterized by the third text, and use a time series analysis algorithm to determine whether the events in the category to be processed occur periodically. Among them, the time series analysis algorithm can be: Time-series Dense Encoder (TiDE) for long-term time series prediction, or Long Short-Term Memory (LSTM), etc.
[0305] If the events in the category to be processed occur periodically, the occurrence period (such as year, month, week, or day, etc.) of the events in the category to be processed can be determined.
[0306] For example, each event characterized by the third text included in a category to be processed is an event related to the failure of the water supply pipeline in a certain area, and the occurrence time of the events in this category to be processed is concentrated in December every year. Correspondingly, the electronic device can use a time series analysis algorithm to obtain that the occurrence period of the events in the category to be processed is one year. Subsequently, the managers of the relevant platform can inform the maintainer of the water supply pipeline to overhaul and maintain the water supply pipeline before December every year, replace the aging water pipes, and stabilize the water supply pressure.
[0307] For another example, all the events of the third text representations included in a category to be processed are events related to the insufficient number of service personnel in a certain regional traffic management department, and the occurrence times of the events in this category to be processed are concentrated on Mondays every week. Correspondingly, the electronic device can use a time series analysis algorithm to obtain that the occurrence period of the events in the category to be processed is one week. Subsequently, the manager of the relevant platform can inform this traffic management department to increase the on-duty staff on Mondays to provide corresponding services to the visitors in a timely manner.
[0308] Based on the above processing, the electronic device can perform periodic analysis on the events in the category to be processed, evaluate the occurrence frequency and trend of the events in this category to be processed, determine the occurrence pattern and specific time period of the events, and provide a basis for preventing and processing the events of this category. Subsequently, the relevant personnel of the events can, according to the analysis result of the occurrence period of the events in this category, achieve the active prevention of the occurrence of the events, reduce the occurrence frequency of the events, reduce the influence range of the events, improve the efficiency of event processing, and realize the transformation of the events from "occurring first and then being processed" to "being processed before occurring".
[0309] Based on the same inventive concept, an embodiment of the present application provides an event recognition device. Refer to Figure 6 , Figure 6 which is a structural diagram of an event recognition device provided by an embodiment of the present application. The device includes:
[0310] A first text acquisition module 601, configured to acquire a plurality of texts containing event information of events as first texts;
[0311] A second text determination module 602, configured to determine, based on a preset clustering algorithm, texts whose represented events belong to the category to be processed from each first text as second texts;
[0312] A theme view generation module 603, configured to input each second text and a first prompt word into a large language model to obtain a first theme view of the texts of the category to be processed; wherein, the first prompt word is used to instruct the large language model to generate the theme view of the input text;
[0313] A matching module 604, configured to use the large language model to determine texts matching the first theme view from each second text as third texts representing events to be processed.
[0314] In some embodiments, the second text determination module 602 includes:
[0315] An encoding sub-module, configured to encode each first text to obtain a first embedding vector of each first text;
[0316] A clustering sub-module, configured to cluster the first embedding vectors based on a preset clustering algorithm to obtain at least one vector cluster;
[0317] A vector cluster selection sub-module, configured to select one vector cluster belonging to the category to be processed from the at least one obtained vector cluster;
[0318] A second text determination sub-module, configured to use the first text represented by the first embedding vectors included in the selected vector cluster as the second text.
[0319] In some embodiments, the apparatus further includes:
[0320] A parameter acquisition module, configured to input the average length, standard deviation of each first text, and a third prompt word into a large language model before encoding each first text to obtain a first parameter for dimensionality reduction processing and a second parameter for clustering;
[0321] Wherein, the first parameter includes at least one of the following: the dimension of the embedding vectors after dimensionality reduction, the minimum distance between the embedding vectors after dimensionality reduction, and the number of adjacent points when constructing a neighborhood after dimensionality reduction; the second parameter includes at least one of the following: the minimum number of vectors included in a vector cluster, the minimum number of vectors in the neighborhood of the center of a vector cluster; the third prompt word is used to: instruct the large language model to generate parameters for dimensionality reduction processing and parameters for clustering based on the average length and standard deviation of the input text;
[0322] The encoding sub-module is specifically configured to:
[0323] Encode each first text and perform dimensionality reduction processing on the encoding result according to the first parameter to obtain the first embedding vectors of each first text;
[0324] The clustering sub-module is specifically configured to:
[0325] Cluster the first embedding vectors based on a preset clustering algorithm according to the second parameter to obtain at least one vector cluster.
[0326] In some embodiments, the matching module 604 is specifically configured to:
[0327] For each second text, input the second text and the first prompt word into the large language model to obtain the second topic view of each second text;
[0328] Calculate the similarity between the first topic view and the second topic view of each second text;
[0329] If the calculated similarity is greater than the preset similarity threshold, then determine the second text as the third text representing the event to be processed.
[0330] In some embodiments, the matching module 604 is specifically configured to:
[0331] For each second text, input the second text, the first theme view, and the fourth prompt word into the large language model to obtain the matching degree of the second text; wherein, the fourth prompt word is used to instruct the large language model to generate the matching degree between the text and the theme view;
[0332] If the matching degree of the second text is greater than the preset matching degree threshold, then determine the second text as the third text representing the event to be processed;
[0333] And / or,
[0334] The first text acquisition module 601 is specifically configured to:
[0335] Acquire multiple texts containing event information of events;
[0336] Preprocess each acquired text respectively to obtain multiple first texts; wherein, the preprocessing includes at least one of the following: denoising, stop word filtering, and stemming;
[0337] And / or,
[0338] The device further includes:
[0339] A period analysis module, configured to obtain an analysis result of the occurrence period of the events in the category to be processed by using a time series analysis algorithm based on the occurrence time of the events represented by the third text.
[0340] In some embodiments, the device further includes:
[0341] A title generation module, configured to input the third text and the second prompt word into the large language model to obtain a title of the text of the category to be processed; wherein, the second prompt word is used to instruct the large language model to generate a title of the input text;
[0342] And / or,
[0343] An occurrence cause acquisition module, configured to, after determining, by using the large language model, a text that matches the first theme view from each second text as the third text representing the event to be processed, input the third text and the fifth prompt word into the large language model to obtain the occurrence cause of the events in the category to be processed; wherein, the fifth prompt word is used to instruct the large language model to generate the occurrence cause of the events represented by the input text;
[0344] and / or
[0345] A risk index acquisition module, after determining, by using the large language model, texts matching the first theme view from each second text as third texts representing events to be processed, inputs the third texts and a sixth prompt word into the large language model to obtain risk indexes of events in the category of events to be processed; wherein, the sixth prompt word is used to instruct the large language model to generate risk indexes of events represented by the input texts, and the risk index of an event includes at least one of the following: the amount involved in the event, the number of people involved in the event, and the regional scope affected by the event;
[0346] and / or
[0347] A risk level acquisition module, after determining, by using the large language model, texts matching the first theme view from each second text as third texts representing events to be processed, inputs the third texts and a seventh prompt word into the large language model to obtain risk levels of events in the category of events to be processed; wherein, the seventh prompt word is used to instruct the large language model to generate risk levels of events represented by the input texts, and the risk level of an event is determined based on the risk index of the event.
[0348] An embodiment of the present invention further provides an electronic device, as Figure 7 shown, including a processor 701, a communication interface 702, a memory 703, and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704.
[0349] The memory 703 is used to store a computer program;
[0350] When the processor 701 executes the program stored on the memory 703, it implements the steps of any of the above event recognition methods.
[0351] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0352] The communication interface is used for communication between the above electronic device and other devices.
[0353] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0354] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0355] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above event recognition methods are implemented.
[0356] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions, which when running on a computer, causes the computer to execute any of the event recognition methods in the above embodiments.
[0357] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a solid-state disk (SSD), etc.
[0358] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.
[0359] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the embodiments of the apparatus, electronic device, and computer-readable storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0360] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included within the protection scope of the present invention.
Claims
1. An event recognition method, characterized in that, The method includes: Obtaining a plurality of texts containing event information of events as first texts; Based on a preset clustering algorithm, determining texts whose represented events belong to the category to be processed from each first text as second texts; Inputting each second text and a first prompt word into a large language model to obtain a first topic view of the text of the category to be processed; wherein, the first prompt word is used to instruct the large language model to generate a topic view of the input text; Using the large language model, determining texts matching the first topic view from each second text as third texts representing the events to be processed; The determining, based on a preset clustering algorithm, texts whose represented events belong to the category to be processed from each first text as second texts includes: Encoding each first text to obtain a first embedding vector of each first text; Based on a preset clustering algorithm, clustering the first embedding vectors to obtain at least one vector cluster; From the obtained at least one vector cluster, selecting a vector cluster belonging to the category to be processed; Taking the first text represented by the first embedding vector included in the selected vector cluster as the second text; Before encoding each first text to obtain a first embedding vector of each first text, the method further includes: Inputting the average length, standard deviation of each first text, and a third prompt word into a large language model to obtain a first parameter for dimensionality reduction processing and a second parameter for clustering; Wherein, the first parameter includes at least one of the following: the dimension of the embedding vector after dimensionality reduction, the minimum distance between each embedding vector after dimensionality reduction, and the number of adjacent points when constructing a neighborhood after dimensionality reduction; the second parameter includes at least one of the following: the minimum number of vectors included in a vector cluster, the minimum number of vectors in the neighborhood of the center of a vector cluster; the third prompt word is used to: instruct the large language model to generate a parameter for dimensionality reduction processing and a parameter for clustering based on the average length and standard deviation of the input text; The encoding each first text to obtain a first embedding vector of each first text includes: Encoding each first text and performing dimensionality reduction processing on the encoding result according to the first parameter to obtain a first embedding vector of each first text; The clustering the first embedding vectors based on a preset clustering algorithm to obtain at least one vector cluster includes: Clustering the first embedding vectors based on a preset clustering algorithm according to the second parameter to obtain at least one vector cluster.
2. The method according to claim 1, characterized in that, The using the large language model to determine texts matching the first topic view from each second text as third texts representing the events to be processed includes: For each second text, inputting the second text and the first prompt word into the large language model to obtain a second topic view of each second text; Calculating the similarity between the first topic view and the second topic view of each second text; If the calculated similarity is greater than a preset similarity threshold, determining the second text as the third text representing the events to be processed.
3. The method according to claim 1, characterized in that Using the large language model to determine, from each second text, a text that matches the first theme view as a third text representing the event to be processed, including: For each second text, input the second text, the first theme view, and a fourth prompt into the large language model to obtain the matching degree of the second text; wherein, the fourth prompt is used to indicate the matching degree between the text generated by the large language model and the theme view; if the matching degree of the second text is greater than a preset matching degree threshold, then determine the second text as the third text representing the event to be processed; and / or The obtaining of multiple texts containing event information of events as the first text includes: Obtaining multiple texts containing event information of events; respectively preprocessing each obtained text to obtain multiple first texts; wherein, the preprocessing includes at least one of the following: denoising, stop word filtering, and stemming; and / or The method further includes: Based on the occurrence time of the event represented by the third text, using a time series analysis algorithm to obtain an analysis result of the occurrence period of the events in the category to be processed.
4. The method according to claim 1, characterized in that, After using the large language model to determine, from each second text, a text that matches the first theme view as a third text representing the event to be processed, the method further includes: Inputting the third text and a second prompt into the large language model to obtain the title of the text in the category to be processed; wherein, the second prompt is used to indicate the large language model to generate the title of the input text; and / or Inputting the third text and a fifth prompt into the large language model to obtain the cause of the event in the category to be processed; wherein, the fifth prompt is used to indicate the large language model to generate the cause of the event represented by the input text; and / or Inputting the third text and a sixth prompt into the large language model to obtain the risk indicator of the event in the category to be processed; wherein, the sixth prompt is used to indicate the large language model to generate the risk indicator of the event represented by the input text, and the risk indicator of an event includes at least one of the following: the amount involved in the event, the number of people involved in the event, the regional scope affected by the event; and / or Inputting the third text and a seventh prompt into the large language model to obtain the risk level of the event in the category to be processed; wherein, the seventh prompt is used to indicate the large language model to generate the risk level of the event represented by the input text, and the risk level of an event is determined based on the risk indicator of the event.
5. An event recognition device, characterized in that, The device includes: A first text acquisition module, configured to acquire multiple texts containing event information of events as the first text; A second text determination module, configured to determine, based on a preset clustering algorithm, texts in which the represented events belong to the category to be processed from each first text as the second text; A theme view generation module for inputting each second text and a first prompt into a large language model to obtain a first theme view of the text of the category to be processed; wherein, the first prompt is used to instruct the large language model to generate a theme view of the input text. A matching module for using the large language model to determine, from each second text, a text that matches the first theme view as a third text representing the event to be processed. The second text determination module includes: An encoding sub-module for encoding each first text to obtain a first embedding vector of each first text. A clustering sub-module for clustering the first embedding vectors based on a preset clustering algorithm to obtain at least one vector cluster. A vector cluster selection sub-module for selecting, from the obtained at least one vector cluster, a vector cluster belonging to the category to be processed. A second text determination sub-module for using the first text represented by the first embedding vector included in the selected vector cluster as the second text. The apparatus further includes: A parameter acquisition module for, before encoding each first text to obtain a first embedding vector of each first text, inputting the average length, standard deviation of each first text, and a third prompt into the large language model to obtain a first parameter for dimensionality reduction processing and a second parameter for clustering. Wherein, the first parameter includes at least one of the following: the dimension of the embedding vector after dimensionality reduction, the minimum distance between each embedding vector after dimensionality reduction, and the number of adjacent points when constructing a neighborhood after dimensionality reduction; the second parameter includes at least one of the following: the minimum number of vectors included in a vector cluster, the minimum number of vectors in the neighborhood of the center of a vector cluster; the third prompt is used to: instruct the large language model to generate a parameter for dimensionality reduction processing and a parameter for clustering based on the average length and standard deviation of the input text. The encoding sub-module is specifically used for: Encoding each first text and performing dimensionality reduction processing on the encoding result according to the first parameter to obtain a first embedding vector of each first text. The clustering sub-module is specifically used for: Clustering the first embedding vectors based on a preset clustering algorithm according to the second parameter to obtain at least one vector cluster.
6. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete mutual communication through the communication bus. The memory is used for storing a computer program. The processor is used for, when executing the program stored on the memory, implementing the method according to any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1-4 is implemented.
8. A computer program product, characterized in that, When the computer program product runs on a computer, the computer is caused to execute the method according to any one of claims 1-4.
Citation Information
Patent Citations
Public health safety event detection and event set construction method and system
CN113449101A
Information processing method and device, equipment, storage medium and program product
CN118708676A