Social governance event similarity matching method and system oriented to multiple data centers
By representing the characteristics of the relationship between event elements and personnel in the text, the weighted similarity calculation method is used to solve the problems of low efficiency and poor accuracy in traditional event retrieval, efficient and accurate matching of event similarity is achieved, and the search quality of social governance is improved.
Patent Information
- Application Number
- CN202411734870.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-07-29
AI Technical Summary
Traditional event search methods consume a lot of time and energy in the field of social governance, and the effect is not satisfactory, making it difficult to accurately identify and compare the similarity information of events in different periods, affecting the efficiency and quality of search work.
By representing the characteristic relationship between event elements and personnel in the text, the similarity is calculated by weighted average or weighted summing, and combined with the processing of structured and unstructured text, the comprehensive similarity value is obtained and the relevant text with high similarity is determined.
It has achieved more accurate identification and comparison of similarity information of events in different periods, improved the efficiency and quality of retrieval work, and met the needs of social governance.
Smart Images

Figure CN120386852A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and specifically relates to a social governance event similarity matching method and system for multiple data centers. Background Art
[0002] In recent years, with the advent of the big data era, the field of social governance has also paid more and more attention to data-driven decision-making and handling. For example, when handling an event, it is often necessary to find historical events similar to the current event in order to better understand the facts of the event and applicable laws, etc.
[0003] However, traditional event retrieval methods often require a large amount of time and effort in practical applications, and their effects are often unsatisfactory. These methods usually rely on keyword matching or rule-based search strategies, which makes them unable to cope when dealing with complex events. For example, when the retrieved event involves multiple contexts or relevant information, simple keyword matching is difficult to capture the true meaning and importance of the event. In addition, traditional methods have insufficient understanding of the grammar structure and context of the text, resulting in a large amount of irrelevant information in the retrieval results, and users need to spend extra time to screen.
[0004] Therefore, how to obtain a more efficient, accurate and intelligent event retrieval method, so as to more accurately identify and compare the similarity information between events in different periods and the current event, improve the efficiency and quality of the retrieval work, meet the needs of social governance construction, and promote the application and development of artificial intelligence technology in the field of social governance, is an urgent problem to be solved currently. Summary of the Invention
[0005] Embodiments of this application provide a social governance event similarity matching method and system for multiple data centers, so as to more accurately identify and compare the similarity information between events in different periods and the current event, and improve the efficiency and quality of the retrieval work.
[0006] To achieve the above object, this application adopts the following technical solutions:
[0007] In a first aspect, the present application provides a social governance event similarity matching method for multiple data centers. The method includes: extracting event elements from the text for which similar events need to be searched, obtaining the personnel-location-correlation relationship of the text, and representing the event elements and the personnel-location-correlation relationship of the text using features; obtaining multiple relevant texts, extracting event elements from the multiple relevant texts, obtaining the personnel-location-correlation relationship of the multiple relevant texts, and representing the event elements and the personnel-location-correlation relationship of the multiple relevant texts using features; calculating the similarity of the event elements and the personnel-location-correlation relationship between the text and the multiple relevant texts one by one based on the feature representations of the event elements and the personnel-location-correlation relationship of the text and the multiple relevant texts; setting the similarity weights of the event elements and the personnel-location-correlation relationship of the text and the multiple relevant texts based on different text situations and requirements, and performing weighted averaging or weighted summation on the similarities of the event elements and the personnel-location-correlation relationship of the text and the multiple relevant texts one by one to obtain a comprehensive similarity value; comparing the comprehensive similarity values between the text and the multiple relevant texts, and determining at least one relevant text with a high comprehensive similarity value.
[0008] In a possible design, the method in the first aspect further includes that the event elements include at least one of the following: event time, event location, event participants, event type, event result.
[0009] In a possible design, the method in the first aspect further includes that obtaining the personnel-location-correlation relationship of the multiple relevant texts includes: dividing each of the multiple relevant texts into two parts: a structured text part and an unstructured text part, where the structured text part is in accordance with certain rules and formats, and the unstructured text part is not in accordance with fixed rules and formats; processing the structured text part based on a rule-based method to obtain basic information about people and locations; processing the unstructured text part based on a deep learning method to obtain event relation triples; and performing knowledge fusion on the basic information about people and locations and the event relation triples to obtain the personnel-location-correlation relationship of each of the multiple relevant texts.
[0010] A possible design solution. The method of the first aspect further includes a rule-based method for processing the structured text part to obtain basic information about people and places, and a deep learning-based method for processing the unstructured text part to obtain event relation triples; performing knowledge fusion on the basic information about people and places and the event relation triples to obtain the person-location-event association relationships of each of the multiple relevant texts, including: finding out the common expression ways of the basic information about people and places according to the grammar structure rules of the structured text part, formulating extraction rules, and obtaining the basic information about people and places; cleaning and preprocessing the unstructured text part, splitting the unstructured text into a sentence list by sentences; performing entity recognition on the sentences in the sentence list based on a deep learning model to obtain an entity list; based on predefined entity relationships, using an entity relationship extraction model to analyze the entity list item by item, extracting the entity relationships, and further obtaining event relation triples; integrating and aligning the same or related entity information in the basic information about people and places and the event relation triples to obtain the person-location-event association relationships of each of the multiple relevant texts.
[0011] A possible design solution. The method of the first aspect further includes representing the event elements and person-location-event association relationships of multiple relevant texts in features, including: converting the event elements of multiple relevant texts into a semantic-acquirable representation form, and converting the person-location-event association relationships of multiple relevant texts into a representation form for calculating similarity.
[0012] A possible design solution. The method of the first aspect further includes calculating the similarity of the event elements and person-location-event association relationships between the text and multiple relevant texts one by one based on the feature representations of the text and the event elements and person-location-event association relationships of multiple relevant texts, including: calculating the similarity of the event elements between the text and multiple relevant texts one by one based on the feature representations of the text and the event elements of multiple relevant texts, and the similarity calculation of the event elements is based on semantic similarity calculation; calculating the similarity of the person-location-event association relationships between the text and multiple relevant texts one by one based on the feature representations of the text and the person-location-event association relationships of multiple relevant texts; wherein, the similarity calculation method is any one of Euclidean distance, cosine similarity, Jaccard similarity, and Hamming distance.
[0013] A possible design solution. The method of the first aspect further includes comparing the comprehensive similarity values between the text and multiple relevant texts respectively, and selecting the relevant text with a high comprehensive similarity value, including: sorting the comprehensive similarity values between the text and multiple relevant texts respectively from high to low, and selecting the relevant text with a high comprehensive similarity value; or, presetting a threshold for the comprehensive similarity, determining and selecting at least one relevant text whose comprehensive similarity value between the text and multiple relevant texts is greater than the threshold.
[0014] Second aspect, the present application provides a social governance event similarity matching system for multiple data centers, the system includes: a relevant text acquisition module, configured to acquire multiple relevant texts of the text for which similar events need to be found; an event element extraction module, configured to extract event elements from the text, and to extract event elements from multiple relevant texts; a personnel-location association relationship acquisition module, configured to acquire the personnel-location association relationship of the text, and to acquire the personnel-location association relationship of multiple relevant texts; a feature representation module, configured to represent the event elements and the personnel-location association relationship of the text with features, and to represent the event elements and the personnel-location association relationship of multiple relevant texts with features; a similarity calculation module, configured to calculate the similarity of the event elements and the personnel-location association relationship between the text and multiple relevant texts one by one based on the feature representations of the event elements and the personnel-location association relationship of the text and multiple relevant texts; configured to set the similarity weights of the event elements and the personnel-location association relationship of the text and multiple relevant texts based on different text situations and requirements, and to perform weighted average or weighted summation on the similarities of the event elements and the personnel-location association relationship of the text and multiple relevant texts one by one to obtain a comprehensive similarity value; a determination module, configured to compare the comprehensive similarity values between the text and multiple relevant texts, and to determine at least one relevant text with a high comprehensive similarity value.
[0015] A possible design solution, the method of the second aspect further includes, the event elements include at least one of the following: event time, event location, event participants, event type, event result.
[0016] A possible design solution, the method of the second aspect further includes, a personnel-location association relationship acquisition module, configured to acquire the personnel-location association relationship of multiple relevant texts, including: a personnel-location association relationship acquisition module, configured to divide each of the multiple relevant texts into two parts: a structured text part and an unstructured text part, wherein, the structured text part is in accordance with certain rules and formats, and the unstructured text part is not in accordance with fixed rules and formats; configured to process the structured text part based on a rule-based method to obtain basic information of people and locations; configured to process the unstructured text part based on a deep learning method to obtain event relationship triples; configured to perform knowledge fusion on the basic information of people and locations and the event relationship triples to obtain the personnel-location association relationship of each of the multiple relevant texts.
[0017] In the embodiments of the present application, for the text that needs to find similar events and the obtained multiple relevant texts, event elements are extracted and the personnel-location association relationship is obtained. The similarity is calculated from two dimensions of event elements and the personnel-location association relationship, and weighted average or weighted summation is performed through the weight settings of different parts to obtain a comprehensive similarity value. By comparing the comprehensive similarity values, at least one relevant text with a high comprehensive similarity value is determined to obtain more comprehensive and accurate similar events. Furthermore, the similarity information between events in different periods and the current event can be more accurately identified and compared, improving the efficiency and quality of the retrieval work.
[0018] Other features and advantages of the present application will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 Flow schematic of the social governance event similarity matching method for multi-data centers provided by the embodiments of the present application Figure 1 ;
[0021] Figure 2 Flow schematic diagram of the social governance event similarity matching method for multi-data centers provided by the embodiments of the present application;;
[0022] Figure 3 Flow schematic of the social governance event similarity matching method for multi-data centers provided by the embodiments of the present application Figure 3 . DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application. At the same time, in the description of the embodiments of the present application, terms such as "first" and "second" are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, "a plurality" means two or more, unless otherwise specifically defined.
[0024] Figure 1 Flow schematic of the social governance event similarity matching method for multiple data centers provided by the embodiments of the present application Figure 1 The social governance event similarity matching method for multiple data centers is applicable to scenarios where, when handling an event, it is necessary to find similar historical events to better understand the facts of the event and applicable laws, etc., or other possible implementation manners, which are not limited herein.
[0025] The process of the social governance event similarity matching method for multiple data centers is as follows:
[0026] Step S101: Extract event elements from the text for which similar events need to be found, obtain the personnel-place-time association relationship of the text, and represent the event elements and personnel-place-time association relationship of the text using features.
[0027] The above text for which similar events need to be found can be understood as the event that needs to be handled currently. The handling of this event requires reference to the handling methods of historical events. A common scenario is that before a judge makes a judicial judgment, a large number of judicial cases often need to be retrieved and consulted for reference. In the following steps, this text for which similar events need to be found is abbreviated as text and will not be elaborated further.
[0028] Step S102: Obtain multiple relevant texts, extract event elements from the multiple relevant texts, obtain the personnel-place-time association relationship of the multiple relevant texts, and represent the event elements and personnel-place-time association relationship of the multiple relevant texts using features.
[0029] The multiple relevant texts retrieved are the relevant texts of the text for which similar events need to be found in the above step S101. In addition, the acquisition method of the multiple relevant texts in the text can be direct acquisition, or can be through API calls. Some websites and platforms provide API interfaces, allowing specific data to be requested and obtained programmatically, or can be database queries, accessing academic databases or open datasets, first obtaining relevant text information through keyword search and screening, etc., which are not limited herein.
[0030] Among them, in step S101 and step S102, extracting event elements from the text for which similar events need to be found, and extracting event elements from the multiple relevant texts. Among them, the event elements include at least one of the following: event time, event location, event participants, event type, event result.
[0031] That is to say, event element extraction abstracts events in the text in three dimensions: people, places, and things, and constructs an element framework for events. In addition, for events in different scenarios, appropriate content such as "key points of handling", "relevant laws and regulations", "basic descriptions", "handling results", "reasons for handling", etc. can be added during the process of event element extraction.
[0032] Obtaining the person-place-thing association relationships of multiple relevant texts in step S102 includes steps S201 - S204, and the specific process is as Figure 2 shown:
[0033] Step S201 divides each of the multiple relevant texts into two parts: a structured text part and an unstructured text part. Among them, the structured text part follows certain rules and formats, and the unstructured text part does not follow fixed rules and formats.
[0034] Among them, structured text usually presents information in the form of tables or fixed formats. For example, person information includes: name: Zhang San, gender: male, age: 30 years old, address: a certain community in Chaoyang District, Beijing, place of birth: Shanghai, educational background: undergraduate, date of birth: May 1993, occupation: software engineer, household registration location: Beijing; place information includes: school: Peking University, park: People's Park, shopping mall: Wangfujing Shopping Center, parking lot: the underground parking lot of a large shopping mall, etc.
[0035] Unstructured text is usually described in natural language, and the information has no fixed format, making it difficult to automatically extract specific features. It can be a typical event description. For example: On September 25, 2023, Zhang San held a lecture on environmental protection in People's Park, attracting many citizens to participate. In this event, Zhang San shared his research results in the field of environmental protection and had in-depth exchanges with the participants. Due to the flexibility and diversity of unstructured text, the processing is often more complex.
[0036] Step S202 processes the structured text part based on a rule-based method to obtain basic person and place information.
[0037] Correspondingly, for the structured text in step S201, according to the grammar structure rules of the structured text part, the common expression ways of basic person and place information can be found, extraction rules can be formulated, and basic person and place information can be obtained. Exemplarily, regular expressions or simple text pattern matching are used to identify time, people, places, etc. Time can be matched with a specific format (such as "YYYY-MM-DD").
[0038] Step S203 processes the unstructured text part based on a deep learning method to obtain event relationship triples.
[0039] Clean and preprocess the unstructured text part, and split the unstructured text into sentences to obtain a sentence list. Cleaning and preprocessing may include: (1) Removing noise: deleting redundant spaces, line breaks, and tab characters. Clearing HTML tags and special characters (such as @, #, etc.). (2) Converting case: converting all text to lowercase to reduce duplicates. (3) Punctuation handling: deleting punctuation in the text or retaining specific punctuation as needed. (4) Spelling correction: identifying and correcting spelling mistakes to improve text quality. (5) Text normalization: handling synonyms and unifying different expressions, for example, treating "car" and "automobile" as the same, and also splitting the unstructured text into sentences and words: splitting the text into sentences and words for subsequent analysis. Or other possible processing methods, which are not specifically restricted here.
[0040] Subsequently, for the sentences in the sentence list obtained from the unstructured text, based on deep learning models such as the trained entity recognition models BERT-CRF, Transformers, Span-Based Models, etc., perform entity recognition to obtain an entity list. The entity categories may be: person names, place names, organization names, dates, times, currencies, events, professional terms, etc. Among them, the entity list includes: entity, entity corresponding category, and the sentence where the entity is located.
[0041] Based on the predefined entity relationships, use the entity relationship extraction model to analyze the entity list item by item, extract the entity relationships, and then obtain the event relationship triples. Among them, the predefined entity relationships may be the relationships that may exist between entities, such as "belong to", "be located in", "participate in", etc. After that, use machine learning models (such as SVM, random forest) or deep learning models (such as LSTM, BERT) to train the entity relationship extraction model, analyze the entity list item by item, extract the entity relationships, and then obtain the event relationship triples. The event relationship triples usually have the format (entity 1, relationship, entity 2), for example: (Li Hua, lives in, Beijing).
[0042] Step S204, perform knowledge fusion on the basic information of people and places and the event relationship triples to obtain the personnel-place associations of each of the multiple relevant texts.
[0043] That is to say, integrate and align the same or relevant entity information in the basic information of people and places and the event relationship triples to obtain the personnel-place associations of each of the multiple relevant texts.
[0044] It should be noted that for obtaining the personnel-place associations of the text in step S101, it can be understood by referring to the above steps S201 - S204 and will not be elaborated here.
[0045] Step S103: Based on the feature representations of the event elements and the person-location-time association relationships of the text and multiple related texts, calculate the similarity of the event elements and the person-location-time association relationships between the text and multiple related texts one by one.
[0046] Convert the event elements of multiple related texts into a semantic-reachable representation form, and convert the person-location-time association relationships of multiple related texts into a similarity-calculable representation form.
[0047] Similarly, convert the event elements of multiple texts into a semantic-reachable representation form, and convert the person-location-time association relationships of multiple texts into a similarity-calculable representation form.
[0048] It can be understood that the event elements can use a word embedding model (such as Word2Vec, BERT, etc.) to convert the extracted event elements into vector representations. For example, participants, time, location, and actions can generate vectors respectively. The person-location-time association relationships can be converted into calculable vector representations, and relationship vectors can be generated by combining the vectors of participants, locations, and events.
[0049] In addition, based on the feature representations of the event elements of the text and multiple related texts, calculate the similarity of the event elements between the text and multiple related texts one by one. The similarity calculation of the event elements is based on semantic similarity calculation; based on the feature representations of the person-location-time association relationships of the text and multiple related texts, calculate the similarity of the person-location-time association relationships between the text and multiple related texts one by one;
[0050] Among them, the similarity calculation method is any one of Euclidean distance, cosine similarity, Jaccard similarity, and Hamming distance.
[0051] Step S104: Based on different text situations and requirements, set the similarity weights of the event elements and the person-location-time association relationships of the text and multiple related texts, and perform weighted average or weighted summation on the similarities of the event elements and the person-location-time association relationships of the text and multiple related texts one by one to obtain a comprehensive similarity value.
[0052] Different texts may have different focuses. Some focus on the participants in the event elements, some focus on the locations in the event elements, some focus on the person-location-time association relationships, etc. According to different considerations, corresponding similarity weights can be selected, and then multiple similarity values are weighted averaged or weighted summed to obtain a comprehensive similarity value.
[0053] In addition, the frequencies of different event elements and person-location-time association relationships can also be considered, and different weight values are set to obtain a comprehensive similarity value.
[0054] Step S105: Compare the comprehensive similarity values between the text and multiple relevant texts, and determine at least one relevant text with a high comprehensive similarity value.
[0055] In a possible implementation, sort the comprehensive similarity values between the text and multiple relevant texts from high to low, and select the relevant text with a high comprehensive similarity value. For example: The comprehensive similarities between the "text" and "relevant text 1", "relevant text 2", "relevant text 3", and "relevant text 4" are 80%, 89%, 76%, and 40% respectively. Then select the "relevant text 2" with the highest comprehensive similarity value for reference and learning.
[0056] In another possible implementation, preset a threshold for the comprehensive similarity, and determine and select at least one relevant text whose comprehensive similarity value between the text and multiple relevant texts is greater than the threshold. For example, preset the threshold for the comprehensive similarity to 80%. The comprehensive similarities between the "text" and "relevant text 1", "relevant text 2", "relevant text 3", and "relevant text 4" are 80%, 89%, 76%, and 40% respectively. Then select the "relevant text 1" and "relevant text 2" whose comprehensive similarity values are greater than 80% for reference and learning.
[0057] In summary, in the embodiment of the present application, for the text that needs to find similar events and the multiple obtained relevant texts, event elements are extracted and the person-location-time association relationship is obtained. Similarity calculation is performed from two dimensions of event elements and person-location-time association relationship, and weighted average or weighted summation is performed through weight settings in different parts to obtain the comprehensive similarity value. By comparing the comprehensive similarity values, at least one relevant text with a high comprehensive similarity value is determined to obtain more comprehensive and accurate similar events. Furthermore, it realizes more accurately identifying and comparing the similarity information between events in different periods and the current event, and improves the efficiency and quality of the retrieval work.
[0058] The above combines Figure 1 - Figure 2 has described in detail the social governance event similarity matching method provided by the embodiment of the present application for multiple data centers. The following combines Figure 3 introduces the specific application scenario of finding relevant texts in the judicial case judgment for the social governance event similarity matching method for multiple data centers.
[0059] Specifically, as shown in Figure 3, the process of the social governance event similarity matching method for multiple data centers is as follows:
[0060] Step S301: Extract event elements from the event text that needs to find similar events, and obtain multiple texts related to the current event text, and also extract event elements from the multiple relevant texts.
[0061] Among them, the event elements include: key points of handling, relevant laws, handling results, reasons for handling, parties, place of case occurrence, place of case submission, etc. For example: the relevant law is Article 264 of the Criminal Law of the People's Republic of China, the handling result is that the suspect was arrested according to law, the suspect is Zhang, the victim is Li, the place of case occurrence is a certain community in Chaoyang District, Beijing, and the place of case submission is the People's Procuratorate of Chaoyang District, Beijing.
[0062] Step S302: Divide the event text and multiple relevant texts into two parts: a structured text part and an unstructured text part.
[0063] Step S303: Based on a rule-based method, process the structured text part to obtain basic information about people and places.
[0064] Exemplarily, the basic information about people and places in the event text includes: Name: Zhang San, Gender: Male, Age: 30 years old, Address: A certain community in Chaoyang District, Beijing, Place of Birth: Shanghai, Educational Attainment: Bachelor's degree, Date of Birth: May 1993, Occupation: Software engineer, Registered Residence: Beijing, etc. The basic information about people and places in relevant text 1 includes: Name: Li Si, Gender: Male, Age: 28 years old, Address: A certain community in Chaoyang District, Beijing, Place of Birth: Jilin, Educational Attainment: Bachelor's degree, Date of Birth: October 1995, Occupation: Software engineer, Registered Residence: Beijing, etc.
[0065] Step S304: Based on a deep learning method, process the unstructured text part to obtain event relation triples.
[0066] Clean and preprocess the unstructured text part. Split the unstructured text into a list of sentences. For the sentences in the list of sentences obtained from the unstructured text, perform entity recognition based on a deep learning model to obtain a list of entities. Based on predefined entity relationships, use an entity relationship extraction model to analyze the list of entities item by item, extract the entity relationships, and then obtain event relation triples. For example: (Zhang San, lives in, Beijing), (Li Si, works as, software engineer), etc.
[0067] Step S305: Perform knowledge fusion on the basic information about people and places and the event relation triples to obtain the personnel-place associations of the event text and multiple relevant texts respectively.
[0068] Integrate and align the same or relevant entity information in the basic information about people and places and the event relation triples, such as: delete duplicate information and retain the most representative descriptions to obtain the personnel-place associations of the text and multiple relevant texts respectively.
[0069] Step S306: Represent the event elements and the personnel-place associations using features.
[0070] Step S307: Based on the feature representations of the event elements and the relationships between people, things, and locations in the event text and multiple related texts, calculate the similarity of the relationships between event elements and between people, things, and locations in the event text and multiple related texts one by one.
[0071] Among them, the calculation results of the similarity of the relationships between people, things, and locations can be shown in Table 1 as follows:
[0072] Event Person Location Matter … Person 0 1 2 Location 1 0 1 Matter 0 0 0 …
[0073] Table 1
[0074] The calculation of the similarity of event elements can be shown in Table 2 as follows:
[0075] Element Similarity Person 5 Location 3 Matter 2 …
[0076] Table 2
[0077] Step S308: Based on different text situations and requirements, set the similarity weights of the event elements and the relationships between people, things, and locations in the event text and multiple related texts, and perform weighted average or weighted summation on the similarities of the event elements and the relationships between people, things, and locations in the event text and multiple related texts one by one to obtain a comprehensive similarity value.
[0078] Among them, the weight value is set by considering that this judicial event focuses on finding texts related to the motive of the crime or others. In addition, the frequencies of different event elements and the relationships between people, things, and locations are also considered. It can be understood with reference to Table 3.
[0079] Event elements, the correlation between people, locations and matters Frequency Weight value Person 5 2 Location 3 1 Matter 2 1 Person - Person 3 3 Person - Location 2 1 Person - Matter 1 1 Location - Matter 2 2 Key theme 2 5 Exception 1 -10
[0080] Table 3
[0081] Step S309: Compare the comprehensive similarity values between the event text and multiple related texts, and determine at least one related text with a high comprehensive similarity value.
[0082] Select according to the comprehensive similarity. For example, select the one with the highest comprehensive similarity value for reference and learning, or select those with a comprehensive similarity value greater than 80% for reference and learning, or select the top three with the highest comprehensive similarity values for reference and learning, etc. This will not be elaborated here.
[0083] The above Figure 1 - Figure 3 has described in detail the social governance event similarity matching method provided by the embodiments of the present application for multiple data centers. The following will detail the social governance event similarity matching system provided by the embodiments of the present application for execution.
[0084] The system specifically includes: a relevant text acquisition module, an event element extraction module, a personnel-location association relationship acquisition module, a feature representation module, a similarity calculation module, and a determination module, as follows.
[0085] The relevant text acquisition module is used to acquire multiple relevant texts of the text for which similar events need to be searched.
[0086] The event element extraction module is used to extract event elements from the text and to extract event elements from multiple relevant texts.
[0087] It should be noted that the event elements include at least one of the following: event time, event location, event participants, event type, event result.
[0088] The personnel-location association relationship acquisition module is used to acquire the personnel-location association relationship of the text and to acquire the personnel-location association relationship of multiple relevant texts.
[0089] It should be noted that the personnel-location association relationship acquisition module is used to divide each of the multiple relevant texts into two parts: a structured text part and an unstructured text part. Among them, the structured text part is in accordance with certain rules and formats, and the unstructured text part is not in accordance with fixed rules and formats; it is used to process the structured text part based on a rule-based method to obtain basic information about people and locations; it is used to process the unstructured text part based on a deep learning method to obtain event relationship triples; it is used to perform knowledge fusion on the basic information about people and locations and the event relationship triples to obtain the personnel-location association relationship of each of the multiple relevant texts. In addition, the steps of the personnel-location association relationship acquisition module for acquiring the personnel-location association relationship of the text can be understood by referring to the acquisition of the personnel-location association relationship of multiple relevant texts and will not be elaborated here.
[0090] The feature representation module is used to represent the event elements and the personnel-location association relationship of the text by features, and to represent the event elements and the personnel-location association relationship of multiple relevant texts by features.
[0091] The similarity calculation module is used to calculate the similarity of the event elements and the personnel-location association relationship between the text and multiple relevant texts one by one based on the feature representation of the event elements and the personnel-location association relationship of the text and multiple relevant texts; it is used to set the similarity weights of the event elements and the personnel-location association relationship between the text and multiple relevant texts according to different text situations and requirements, and to perform weighted average or weighted summation on the similarity of the event elements and the personnel-location association relationship between the text and multiple relevant texts one by one to obtain a comprehensive similarity value.
[0092] The determination module is used to compare the comprehensive similarity values between the text and multiple relevant texts and to determine at least one relevant text with a high comprehensive similarity value.
[0093] In addition, for the specific implementation of the above system, since it is basically similar to the method implementation, the description is relatively simple. For related parts, please refer to the partial description of the method implementation. Moreover, it should be noted that in each module of the system of the present application, the components are logically divided according to the functions to be achieved. However, the present application is not limited thereto, and the components can be re-divided or combined as needed.
[0094] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.
[0095] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, a particular order or a sequential order shown in the drawings is not necessarily required to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0096] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A social governance event similarity matching method for multi-data centers, characterized in that The method includes: extracting event elements from the text for which similar events need to be searched, obtaining the personnel-location association relationship of the text, and representing the event elements and the personnel-location association relationship of the text using features; obtaining multiple relevant texts, extracting event elements from the multiple relevant texts, obtaining the personnel-location association relationships of the multiple relevant texts, and representing the event elements and the personnel-location association relationships of the multiple relevant texts using features; based on the feature representations of the event elements and the personnel-location association relationships of the text and the multiple relevant texts, calculating the similarity of the event elements and the personnel-location association relationships between the text and the multiple relevant texts one by one; based on different text situations and requirements, setting the similarity weights of the event elements and the personnel-location association relationships of the text and the multiple relevant texts, and performing weighted averaging or weighted summation on the similarities of the event elements and the personnel-location association relationships of the text and the multiple relevant texts one by one to obtain a comprehensive similarity value; comparing the comprehensive similarity values between the text and the multiple relevant texts, and determining at least one relevant text with a high comprehensive similarity value.
2. The social governance event similarity matching method for multiple data centers according to claim 1, wherein The event elements include at least one of the following: event time, event location, event participants, event type, event result.
3. The method for matching social governance event similarities for multiple data centers according to claim 1, wherein The obtaining of the personnel-location association relationships of the multiple relevant texts includes: dividing each of the multiple relevant texts into two parts: a structured text part and an unstructured text part, where the structured text part is in a certain rule and format, and the unstructured text part is not in a fixed rule and format; processing the structured text part based on a rule-based method to obtain basic information about people and locations; processing the unstructured text part based on a deep learning method to obtain event relation triples; performing knowledge fusion on the basic information about people and locations and the event relation triples to obtain the personnel-location association relationships of each of the multiple relevant texts.
4. The social governance event similarity matching method for multiple data centers according to claim 3, characterized in that, Processing the structured text part based on a rule-based method to obtain basic information about people and locations, and processing the unstructured text part based on a deep learning method to obtain event relation triples; Performing knowledge fusion on the basic information about people and locations and the event relation triples to obtain the personnel-location association relationships of each of the multiple relevant texts includes: finding common expression ways of basic information about people and locations according to the grammar structure rules of the structured text part, formulating extraction rules, and obtaining basic information about people and locations; cleaning and preprocessing the unstructured text part, and splitting the unstructured text into a sentence list by sentences; performing entity recognition on the sentences in the sentence list based on a deep learning model to obtain an entity list; based on predefined entity relationships, using an entity relationship extraction model to analyze each item in the entity list one by one, extracting entity relationships, and further obtaining event relation triples; integrating and aligning the same or relevant entity information in the basic information about people and locations and the event relation triples to obtain the personnel-location association relationships of each of the multiple relevant texts.
5. The social governance event similarity matching method for multiple data centers according to claim 1, characterized in that The feature representation of the event elements and the personnel-location correlation relationships of the multiple related texts includes: Converting the event elements of the multiple related texts into a representation form that can obtain semantics, and converting the personnel-location correlation relationships of the multiple related texts into a representation form that can calculate similarity.
6. The method for matching social governance event similarities for multiple data centers according to claim 1, wherein Based on the feature representation of the event elements and the personnel-location correlation relationships of the text and the multiple related texts, the similarity calculation of the event elements and the personnel-location correlation relationships between the text and the multiple related texts is performed one by one, including: Based on the feature representation of the event elements of the text and the multiple related texts, the similarity calculation of the event elements between the text and the multiple related texts is performed one by one, and the similarity calculation of the event elements is based on semantic similarity calculation; Based on the feature representation of the personnel-location correlation relationships of the text and the multiple related texts, the similarity calculation of the personnel-location correlation relationships between the text and the multiple related texts is performed one by one; Among them, the similarity calculation method is any one of Euclidean distance, cosine similarity, Jaccard similarity, and Hamming distance.
7. The method for matching social governance event similarities for multiple data centers according to claim 1, wherein Comparing the comprehensive similarity values between the text and the multiple related texts respectively, and selecting the related texts with high comprehensive similarity values, including: Sorting the comprehensive similarity values between the text and the multiple related texts respectively from high to low, and selecting the related texts with high comprehensive similarity values; Or, presetting a threshold for comprehensive similarity, and determining and selecting at least one related text whose comprehensive similarity value between the text and the multiple related texts is greater than the threshold.
8. A social governance event similarity matching system for multiple data centers, characterized in that, The system includes: A related text acquisition module, configured to acquire multiple related texts of the text for which similar events need to be searched; An event element extraction module, configured to extract event elements from the text and to extract event elements from the multiple related texts; A personnel-location correlation relationship acquisition module, configured to acquire the personnel-location correlation relationship of the text and to acquire the personnel-location correlation relationship of the multiple related texts; A feature representation module, configured to represent the event elements and the personnel-location correlation relationship of the text by features, and to represent the event elements and the personnel-location correlation relationship of the multiple related texts by features; A similarity calculation module, configured to perform the similarity calculation of the event elements and the personnel-location correlation relationships between the text and the multiple related texts one by one based on the feature representation of the event elements and the personnel-location correlation relationships of the text and the multiple related texts; configured to set the similarity weights of the event elements and the personnel-location correlation relationships of the text and the multiple related texts according to different text situations and requirements, and perform weighted average or weighted summation on the similarities of the event elements and the personnel-location correlation relationships between the text and the multiple related texts one by one to obtain a comprehensive similarity value; A determination module, configured to compare the comprehensive similarity values between the text and the multiple related texts, and determine at least one related text with a high comprehensive similarity value.
9. The social governance event similarity matching system for multiple data centers according to claim 8, characterized in that, The event elements include at least one of the following: event time, event location, event participants, event type, and event result.
10. The social governance event similarity matching system for multiple data centers according to claim 8, wherein The person-location-event association relationship acquisition module is used to acquire the person-location-event association relationships of the multiple relevant texts, including: The person-location-event association relationship acquisition module is used to divide each of the multiple relevant texts into two parts: a structured text part and an unstructured text part, where the structured text part is in a certain rule and format, and the unstructured text part is not in a fixed rule and format; It is used to process the structured text part based on a rule-based method to obtain basic information about people and locations; It is used to process the unstructured text part based on a deep learning method to obtain event relationship triples; It is used to perform knowledge fusion on the basic information about people and locations and the event relationship triples to obtain the person-location-event association relationships of the multiple relevant texts respectively.