Similar event duplicate removal method based on cosine similarity
By adding proprietary information to the lexicon and calculating cosine similarity, the problem of similar or repeated event descriptions in events reported by grid workers is solved, and the efficiency and accuracy of event processing are improved.
Patent Information
- Application Number
- CN202510145306.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-16
AI Technical Summary
With the development of urban grid management, the number of events reported by grid workers has increased, resulting in a large number of similar or repeated event descriptions, which has increased the work burden of the handlers and affected the efficiency and accuracy of event processing.
The similar event deduplication method based on cosine similarity is used. By adding proprietary information to the lexicon library, pre-processing the event with the lexicon library, the processed text data and the text data of each event in the event description library are calculated, and whether the event is similar is determined based on the set similarity threshold, and if it is similar, it will be removed.
Effectively remove similar events, improve data quality, improve the efficiency and accuracy of event processing, and reduce the work burden of processors.
Smart Images

Figure CN120012775A_ABST
Abstract
Description
Technical Field
[0001] The invention discloses a similar event deduplication method based on cosine similarity and relates to the field of information technology. Background Art
[0002] With the development of urban grid management, the number of events reported by grid workers is increasing, including a large number of similar or repeated event descriptions. These similar or repeated events not only increase the workload of the processing personnel, but also affect the efficiency and accuracy of event processing. Summary of the invention
[0003] In view of the problems of the prior art, the present invention provides a similar event deduplication method based on cosine similarity, which can effectively remove similar events, improve data quality, and improve the efficiency and accuracy of event processing.
[0004] The specific scheme proposed by the present invention is:
[0005] The present invention provides a method for removing duplicate similar events based on cosine similarity, comprising:
[0006] Step 1: Add proprietary information to the word segmentation database, including the grid name and grid operator name;
[0007] Step 2: Pre-process the events reported by grid workers in combination with the word segmentation database to obtain processed text data;
[0008] Step 3: Calculate the cosine similarity between the processed text data and the text data of each event in the event description library:
[0009] The processed text data and the text data in the event description library are processed by word frequency vectorization.
[0010] Calculate the multidimensional cosine value of the two text data after word frequency vectorization:
[0011] The reported text data is vectorized and represented as {x1,x2,...x n}, x represents the frequency of occurrence of each word in the reported text data,
[0012] The text data in the event description library is vectorized and represented as {y1,y2,…y n}, y represents the frequency of occurrence of each word in the text data of the event description library,
[0013] Using the following formula:
[0014]
[0015] Calculate the cosine similarity and retain the maximum cosine similarity;
[0016] Step 4: According to the set similarity threshold, determine whether the reported event is similar to the existing events in the event description library. If the maximum value of the retained cosine similarity exceeds the similarity threshold, it is judged as a similar event, otherwise it is a dissimilar event;
[0017] Step 5: If they are similar, the reported event is marked as a similar event and removed; otherwise, the text data of the reported event is added to the event description library.
[0018] Furthermore, in the method for deduplicating similar events based on cosine similarity, in step 1, the jieba word segmentation library is selected, a new text file is created, grid names and grid member names are collected, the collected grid names and grid member names are added to the text file, and the load_userdict function is used to load the text file into the jieba word segmentation library to add proprietary information to the word segmentation library.
[0019] Furthermore, in the method for deduplicating similar events based on cosine similarity, in step 2, the events reported by the grid workers are preprocessed in combination with the Jieba word library, including: using a text replacement tool to remove punctuation marks and special characters of the reported events, special characters including spaces and line breaks,
[0020] Then, the lcut function of the jieba word segmentation library is used to segment the reported events. According to the stop word dictionary, the stop words after segmentation are filtered out to obtain text data.
[0021] Furthermore, in the method for deduplicating similar events based on cosine similarity, in step 4, the similarity threshold is set to 0.85 according to the value range [-1, 1] of the cosine similarity. When the maximum retained cosine similarity exceeds 0.85, it is judged as a similar event, otherwise it is a dissimilar event.
[0022] The present invention provides a similar event deduplication device based on cosine similarity, comprising a word segmentation library management module, a text preprocessing module, a similarity calculation module, a judgment module and a similar event removal module.
[0023] The word segmentation database management module adds proprietary information to the word segmentation database, and the proprietary information includes the grid name and the grid member name;
[0024] The text preprocessing module combines the word segmentation library to preprocess the events reported by the grid workers to obtain the processed text data;
[0025] The similarity calculation module calculates the cosine similarity between the processed text data and the text data of each event in the event description library:
[0026] The processed text data and the text data in the event description library are processed by word frequency vectorization.
[0027] Calculate the multidimensional cosine value of the two text data after word frequency vectorization:
[0028] The reported text data is vectorized and represented as {x1,x2,...x n}, x represents the frequency of occurrence of each word in the reported text data,
[0029] The text data in the event description library is vectorized and represented as {y1,y2,…y n}, y represents the frequency of occurrence of each word in the text data of the event description library,
[0030] Using the following formula:
[0031]
[0032] Calculate the cosine similarity and retain the maximum cosine similarity;
[0033] The judgment module judges whether the reported event is similar to the existing events in the event description library according to the set similarity threshold. If the maximum value of the retained cosine similarity exceeds the similarity threshold, it is judged as a similar event, otherwise it is a dissimilar event;
[0034] If they are similar, the similar event removal module marks the reported event as a similar event and removes it; otherwise, the similar event removal module adds the text data of the reported event to the event description library.
[0035] Furthermore, the word segmentation library management module of the similar event deduplication device based on cosine similarity selects the jieba word segmentation library, creates a new text file, collects grid names and grid member names, adds the collected grid names and grid member names to the text file, and uses the load_userdict function to load the text file into the jieba word segmentation library to add proprietary information to the word segmentation library.
[0036] Furthermore, the text preprocessing module of the similar event deduplication device based on cosine similarity combines the Jieba word segmentation library to preprocess the events reported by the grid workers, and uses a text replacement tool to remove punctuation marks and special characters from the reported events, including spaces and line breaks.
[0037] Then, the lcut function of the jieba word segmentation library is used to segment the reported events. According to the stop word dictionary, the stop words after segmentation are filtered out to obtain text data.
[0038] Furthermore, the judgment module of the similar event deduplication device based on cosine similarity sets the similarity threshold to 0.85 according to the value range [-1,1] of cosine similarity, and judges it as a similar event when the retained maximum cosine similarity exceeds 0.85, otherwise it is a dissimilar event.
[0039] The benefits of the present invention are:
[0040] Collect grid management related information, organize proprietary information, add proprietary information to the word segmentation database, and build a word segmentation database containing proprietary information, which can ensure the accuracy and completeness of proprietary information in the word segmentation process and improve the accuracy of text word segmentation;
[0041] Preprocessing the event descriptions reported by grid workers is crucial for extracting key information from the text and optimizing text representation. It can ensure that high-quality text data can be obtained when processing the event descriptions reported by grid workers, providing strong support for subsequent event identification and classification.
[0042] Calculating the cosine similarity of the text can ensure that the similarity between the processed text data and the text data in the event description library can be calculated accurately and efficiently, thereby determining whether the reported event description is similar to the existing event description;
[0043] Marking and removing similar events and adding new event descriptions to the event description library can ensure that when processing events reported by grid workers, they can be accurately classified according to preset rules, avoid duplicate processing, and improve the efficiency and accuracy of event processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a schematic flow chart of the method of the present invention. DETAILED DESCRIPTION
[0045] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.
[0046] Example 1
[0047] The present invention provides a method for removing duplicate similar events based on cosine similarity, comprising:
[0048] Step 1: Add proprietary information to the word segmentation library, including grid names and grid member names. Unique information such as grid names may appear in event descriptions. The existing word segmentation library does not contain such information, which affects the word segmentation results. Therefore, the grid name and other information in the grid should be extracted in advance and added to the word segmentation library. You can select the jieba word segmentation library, create a new text file, collect grid names and grid member names, add the collected grid names and grid member names to the text file, use the load_userdict function to load the text file into the jieba word segmentation library, and add proprietary information to the word segmentation library.
[0049] Step 2: Pre-process the events reported by grid workers in combination with the word segmentation database to obtain processed text data.
[0050] In step 2, the events reported by the grid workers can be pre-processed in combination with the Jieba word library, including: using text replacement tools to remove punctuation marks and special characters in the reported events, special characters include spaces and line breaks,
[0051] Then, the lcut function of the jieba word segmentation library is used to segment the reported events. According to the stop word dictionary, the stop words after segmentation are filtered out to obtain text data.
[0052] Step 3: Calculate the cosine similarity between the processed text data and the text data of each event in the event description library:
[0053] The processed text data and the text data in the event description library are processed by word frequency vectorization.
[0054] Calculate the multidimensional cosine value of the two text data after word frequency vectorization:
[0055] The reported text data is vectorized and represented as {x1,x2,...x n}, x represents the frequency of occurrence of each word in the reported text data,
[0056] The text data in the event description library is vectorized and represented as {y1,y2,…y n}, y represents the frequency of occurrence of each word in the text data of the event description library,
[0057] Using the following formula:
[0058]
[0059] Calculate the cosine similarity and keep the maximum cosine similarity.
[0060] The word frequency vectorization process is performed. The specific operation is as follows: the text data is obtained into a series of separated word sequences, the frequency of occurrence of all words is counted, and a vector is constructed based on the statistical results, where each dimension represents a word, and the value of each dimension represents the number of times the word appears in the text. For example, the two sentences of the text data are "I eat watermelon" and "I drink water". The first text data is ['I', 'eat', 'watermelon'] after word segmentation, and the second text data is ['I', 'drink', 'water']. All words are ['I', 'eat', 'watermelon', 'drink', 'water']. The first text is vectorized to [1,1,1,0,0], because 'I', 'eat', 'watermelon' each appear once, and 'drink', 'water' appear 0 times. The second text is vectorized to [1,0,0,1,1]. Substitute into the formula to calculate the cosine similarity.
[0061] The reported text data is calculated for each text data in the event description library in turn, and the maximum value of the similarity is retained.
[0062] Step 4: According to the set similarity threshold, determine whether the reported event is similar to the existing events in the event description library. If the maximum value of the retained cosine similarity exceeds the similarity threshold, it is judged as a similar event, otherwise it is a dissimilar event.
[0063] According to the value range of cosine similarity [-1,1], 1 means that the two vectors are completely similar, 0 means that the two vectors are unrelated, and -1 means that the two vectors are completely opposite. Therefore, the closer the value is to 1, the more similar the event descriptions are. The similarity threshold is set to 0.85. When the maximum value of the retained cosine similarity exceeds 0.85, it is judged as a similar event, otherwise it is a dissimilar event.
[0064] Step 5: If they are similar, the reported event is marked as a similar event and removed; otherwise, the text data of the reported event is added to the event description library.
[0065] Further, in step 1 of the method for removing duplicate similar events based on cosine similarity, further, in the method for removing duplicate similar events based on cosine similarity
[0066] Further, in the method for deduplicating similar events based on cosine similarity, in step 4 of embodiment 2
[0067] The present invention provides a similar event deduplication device based on cosine similarity, comprising a word segmentation library management module, a text preprocessing module, a similarity calculation module, a judgment module and a similar event removal module.
[0068] The word segmentation database management module adds proprietary information to the word segmentation database, and the proprietary information includes the grid name and the grid member name;
[0069] The text preprocessing module combines the word segmentation library to preprocess the events reported by the grid workers to obtain the processed text data;
[0070] The similarity calculation module calculates the cosine similarity between the processed text data and the text data of each event in the event description library:
[0071] The processed text data and the text data in the event description library are processed by word frequency vectorization.
[0072] Calculate the multidimensional cosine value of the two text data after word frequency vectorization:
[0073] The reported text data is vectorized and represented as {x1,x2,...x n}, x represents the frequency of occurrence of each word in the reported text data,
[0074] The text data in the event description library is vectorized and represented as {y1,y2,…y n}, y represents the frequency of occurrence of each word in the text data of the event description library,
[0075] Using the following formula:
[0076]
[0077] Calculate the cosine similarity and retain the maximum cosine similarity;
[0078] The judgment module judges whether the reported event is similar to the existing events in the event description library according to the set similarity threshold. If the maximum value of the retained cosine similarity exceeds the similarity threshold, it is judged as a similar event, otherwise it is a dissimilar event;
[0079] If they are similar, the similar event removal module marks the reported event as a similar event and removes it; otherwise, the similar event removal module adds the text data of the reported event to the event description library.
[0080] As the information interaction and execution process between the modules in the above-mentioned device are based on the same concept as the embodiment of the method of the present invention, the specific contents can be found in the description of the embodiment of the method of the present invention and will not be repeated here.
[0081] Similarly, the device of the present invention collects grid management related information, organizes proprietary information, adds proprietary information to the word segmentation library, and constructs a word segmentation library containing proprietary information, which can ensure the accuracy and completeness of proprietary information in the word segmentation process and improve the accuracy of text word segmentation;
[0082] Preprocessing the event descriptions reported by grid workers is crucial for extracting key information from the text and optimizing text representation. It can ensure that high-quality text data can be obtained when processing the event descriptions reported by grid workers, providing strong support for subsequent event identification and classification.
[0083] Calculating the cosine similarity of the text can ensure that the similarity between the processed text data and the text data in the event description library can be calculated accurately and efficiently, thereby determining whether the reported event description is similar to the existing event description;
[0084] Marking and removing similar events and adding new event descriptions to the event description library can ensure that when processing events reported by grid workers, they can be accurately classified according to preset rules, avoid duplicate processing, and improve the efficiency and accuracy of event processing.
[0085] It should be noted that not all steps and modules in the above-mentioned processes and device structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The system structure described in the above-mentioned embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or some components in multiple independent devices may be implemented together.
[0086] The above-described embodiments are only preferred embodiments for fully illustrating the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or changes made by those skilled in the art based on the present invention are within the protection scope of the present invention. The protection scope of the present invention shall be subject to the claims.
Claims
1. A method for removing duplicate similar events based on cosine similarity, characterized by include: Step 1: Add proprietary information to the word segmentation database, including the grid name and grid operator name; Step 2: Pre-process the events reported by grid workers in combination with the word segmentation database to obtain processed text data; Step 3: Calculate the cosine similarity between the processed text data and the text data of each event in the event description library: The processed text data and the text data in the event description library are processed by word frequency vectorization. Calculate the multidimensional cosine value of the two text data after word frequency vectorization: The reported text data is vectorized and represented as {x1,x2,...x n }, x represents the frequency of occurrence of each word in the reported text data, The text data in the event description library is vectorized and represented as {y1,y2,…y n }, y represents the frequency of occurrence of each word in the text data of the event description library, Using the following formula: Calculate the cosine similarity and retain the maximum cosine similarity; Step 4: According to the set similarity threshold, determine whether the reported event is similar to the existing events in the event description library. If the maximum value of the retained cosine similarity exceeds the similarity threshold, it is judged as a similar event, otherwise it is a dissimilar event; Step 5: If they are similar, the reported event is marked as a similar event and removed; otherwise, the text data of the reported event is added to the event description library.
2. The method for removing duplicate similar events based on cosine similarity according to claim 1, characterized in that In step 1, select the jieba word segmentation library, create a new text file, collect grid names and grid member names, add the collected grid names and grid member names to the text file, and use the load_userdict function to load the text file into the jieba word segmentation library to add proprietary information to the word segmentation library.
3. The method for removing duplicate similar events based on cosine similarity according to claim 2, characterized in that In step 2, the events reported by the grid workers are preprocessed in combination with the Jieba word library, including: using text replacement tools to remove punctuation marks and special characters in the reported events, special characters include spaces and line breaks, Then, the lcut function of the jieba word segmentation library is used to segment the reported events. According to the stop word dictionary, the stop words after segmentation are filtered out to obtain text data.
4. According to the method for deduplicating similar events based on cosine similarity in claim 1, it is characterized in that in step 4, the similarity threshold is set to 0.85 according to the value range of cosine similarity [-1,1], and when the maximum value of the retained cosine similarity exceeds 0.85, it is judged as a similar event, otherwise it is a dissimilar event.
5. A device for deduplicating similar events based on cosine similarity, characterized in that It includes word segmentation library management module, text preprocessing module, similarity calculation module, judgment module and similar event removal module. The word segmentation database management module adds proprietary information to the word segmentation database, and the proprietary information includes the grid name and the grid member name; The text preprocessing module combines the word segmentation library to preprocess the events reported by the grid workers to obtain the processed text data; The similarity calculation module calculates the cosine similarity between the processed text data and the text data of each event in the event description library: The processed text data and the text data in the event description library are processed by word frequency vectorization. Calculate the multidimensional cosine value of the two text data after word frequency vectorization: The reported text data is vectorized and represented as {x1,x2,...x n }, x represents the frequency of occurrence of each word in the reported text data, The text data in the event description library is vectorized and represented as {y1,y2,…y n }, y represents the frequency of occurrence of each word in the text data of the event description library, Using the following formula: Calculate the cosine similarity and retain the maximum cosine similarity; The judgment module judges whether the reported event is similar to the existing events in the event description library according to the set similarity threshold. If the maximum value of the retained cosine similarity exceeds the similarity threshold, it is judged as a similar event, otherwise it is a dissimilar event; If they are similar, the similar event removal module marks the reported event as a similar event and removes it; otherwise, the similar event removal module adds the text data of the reported event to the event description library.
6. The device for removing duplicate similar events based on cosine similarity according to claim 5 is characterized in that the segmentation The library management module selects the jieba word segmentation library, creates a new text file, collects grid names and grid member names, adds the collected grid names and grid member names to the text file, and uses the load_userdict function to load the text file into the jieba word segmentation library to add proprietary information to the word segmentation library.
7. The device for removing duplicate similar events based on cosine similarity according to claim 6, characterized in that The text preprocessing module combines the Jieba word segmentation library to preprocess the events reported by the grid workers, and uses the text replacement tool to remove punctuation marks and special characters from the reported events. Special characters include spaces and line breaks. Then, the lcut function of the jieba word segmentation library is used to segment the reported events. According to the stop word dictionary, the stop words after segmentation are filtered out to obtain text data.
8. The device for deduplicating similar events based on cosine similarity according to claim 5, characterized in that The judgment module sets the similarity threshold to 0.85 based on the cosine similarity value range [-1,1]. When the retained maximum cosine similarity exceeds 0.85, it is judged as a similar event, otherwise it is a dissimilar event.
Citation Information
Cited By
Method and device for identifying and automatically merging petition documents
CN120877315A
Similar event judgment method and device based on geographic position information conversion and medium
CN121501862A