Telegram Chinese Group Retrieval Method, Device and Equipment for Integrating Multi-Source Data
Through multi-source data fusion and group chat record analysis, the problem of difficulty in searching Telegram Chinese group is solved, comprehensive and accurate group retrieval is achieved, search frequency and update requirements are reduced, and groups related to content can be effectively retrieved.
Patent Information
- Application Number
- CN202211429752.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-15
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-11-15
AI Technical Summary
Telegram Chinese groups are difficult to search, with few and inaccurate search results. The existing methods require frequent traversal and update of the knowledge base, and cannot retrieve groups that are not related to the title but are related to the content.
Through multi-source data fusion technology, Google, Twitter and third-party Telegram group information services are used to obtain the initial group list, combine pinyin search and group chat record analysis, generate a collection of feature words, conduct association association and association search, filter irrelevant groups, and finally generate accurate search results.
A comprehensive and accurate search of Telegram Chinese groups is realized, extensive searches and frequent updates of groups are reduced, and groups that are not related to titles but content can be effectively retrieved.
Smart Images

Figure CN115712738B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information retrieval, and particularly to a method, device and equipment for retrieving Chinese Telegram groups by integrating multi-source data. Background Art
[0002] Telegram is an instant messaging software with a large number of users worldwide. Users can create or join different groups according to their interests and hobbies. Among them, the chat information of public groups can be viewed by any user without joining. However, due to the loose supervision of this software, it contains a large number of groups involving illegal and criminal activities, and illegal activities are still carried out on this software. How to accurately locate groups and timely master illegal and criminal information is of great significance for preventing and cracking down on crimes. However, Telegram only provides an English search function, and it is still difficult to effectively search for Chinese groups related to specific subject words. Some developers accumulate group and title knowledge for Telegram robots and use keywords to match the group titles in the knowledge base to achieve the Chinese group search function. Although this method can achieve the Chinese search function, there are several disadvantages:
[0003] 1) This type of method requires the robot to traverse a large number of groups in advance and accumulate a wide range of knowledge bases;
[0004] 2) The Telegram group title allows for arbitrary changes. If the search accuracy is to be maintained, the knowledge base needs to be traversed and updated frequently;
[0005] 3) When the group title cannot match the search term, but the content of the group is related to the search term, such groups are difficult to be retrieved. Summary of the Invention
[0006] The purpose of the embodiments of the present invention is to provide a method, device and equipment for retrieving Chinese Telegram groups by integrating multi-source data, which focuses on solving problems such as difficult retrieval of Chinese Telegram groups, few retrieval results, and inaccurate retrieval results.
[0007] To achieve the above purpose, the technical solution of the present invention is as follows:
[0008] According to the first aspect of the embodiments of the present disclosure, a method for retrieving Chinese Telegram groups by integrating multi-source data is provided, including:
[0009] Obtain a search term, and perform a search for Chinese Telegram groups on the search term to generate a multi-source fusion group;
[0010] Analyze the set of group chat records corresponding to the multi-source fusion group to obtain a set of feature words, and filter the multi-source fusion group based on the set of feature words to obtain the conforming feature group V0;
[0011] Based on the sharing groups in the set of group chat records corresponding to the conforming feature group V t-1 obtain the associated group R t-1 ; where t represents the number of iteration rounds;
[0012] Perform an associative search on each Telegram Chinese group in the associated group R t-1 to generate the associated associative group L t-1 ;
[0013] Filter the associated group R t-1 and the associated associative group L t-1 based on the feature words to obtain the conforming feature group V t ;
[0014] When the conforming feature group V t is not an empty set, let t = t + 1, and return to perform an associative search on each Telegram Chinese group in the associated group R t-1 to generate the associated associative group L t-1 ;
[0015] When the conforming feature group V t is an empty set, obtain the Telegram Chinese group retrieval result based on the multi-source fusion group and the set of conforming feature groups V; where V = {V0,..., V t-1}.
[0016] Furthermore, the performing a Telegram Chinese group retrieval on the retrieval term to generate a multi-source fusion group includes:
[0017] Use multiple data sources to perform a Telegram Chinese group retrieval on the retrieval term to obtain a multi-source data retrieval group.
[0018] Furthermore, based on the English group retrieval interface provided by Telegram, perform a Telegram Chinese group search on the pinyin of the retrieval term and pinyins similar to the pinyin of the retrieval term to obtain a retrieval term associative group;
[0019] Merge the multi-source data retrieval group and the retrieval term associative group, and perform deduplication to obtain a multi-source fusion group.
[0020] Furthermore, the multiple data sources include: Google data source, Twitter data source, and other third-party Telegram group information retrieval service data sources;
[0021] Performing a Telegram Chinese group search on the search term using multiple data sources to obtain a multi-source data search group, including:
[0022] Performing a directional search for the search term within the scope of telegram.org using a custom search mode to obtain the search results corresponding to the Google data source;
[0023] Using web crawler technology to perform a directional search for the search term in Twitter data and filtering the data containing the Telegram group field to obtain the search results corresponding to the Twitter data source;
[0024] Searching for the search term through the Q&A service of the Telegram robot account in the said other third-party Telegram retrieval service to obtain the search results corresponding to the data source of the other third-party Telegram group information retrieval service;
[0025] Merging the search results corresponding to the Google data source, the search results corresponding to the Twitter data source, and the search results corresponding to the data source of the other third-party Telegram group information retrieval service, and performing deduplication to obtain a multi-source data search group.
[0026] Furthermore, performing a Telegram Chinese group search on the pinyin of the search term and the pinyin similar to the pinyin of the search term based on the English group retrieval interface provided by Telegram to obtain an associated search term group;
[0027] Calculating the pinyin of the search term;
[0028] Generating the pinyin similar to the pinyin of the search term;
[0029] Based on the English group retrieval interface provided by Telegram, and using the pinyin of the search term and the pinyin similar to the pinyin of the search term to retrieve the group username to obtain the first associated search results;
[0030] Based on the English group retrieval interface provided by Telegram, and using the pinyin of the search term to retrieve the group title to obtain the second associated search results;
[0031] Merging the first associated search results and the second associated search results, and performing deduplication to obtain an associated search term group.
[0032] Furthermore, analyzing the set of group chat records corresponding to the multi-source fusion group to obtain a set of feature words, including:
[0033] For the multi-source fusion group, use the word segmentation technology to segment the chat records of each Telegram Chinese group, and generate keyword pairs based on the order of the keywords in the conversation;
[0034] Take the high-frequency words in the word segmentation result as keywords;
[0035] Construct a keyword relationship graph; wherein, the nodes in the keyword relationship graph are the keywords, the edges in the keyword relationship graph represent the associations of the keyword pairs, the weight of the nodes is the number of times the keyword appears, and the weight of the edges is the number of times the keyword pair appears;
[0036] Screen the keywords based on the weights of the nodes to obtain a set of main feature words;
[0037] According to the main feature words and the weights of the edges connecting the main feature words, obtain a set of auxiliary feature words;
[0038] Merge the set of main feature words and the set of auxiliary feature words to obtain a set of feature words.
[0039] Further, perform an associative search on each Telegram Chinese group in the associated group R t-1 to generate an associated associative group L t-1 , including:
[0040] Obtain the group name of each Telegram Chinese group in the associated group R t-1 ;
[0041] Generate approximate group names similar to the group name, and obtain the pinyin of the approximate group names;
[0042] Based on the above, perform a search for Telegram Chinese groups with multiple data sources, and / or perform a Telegram Chinese group search on the pinyin of the approximate group names based on the English group retrieval interface provided by Telegram, to obtain the associated associative group L t-1 .
[0043] According to the second aspect of the embodiments of the present disclosure, there is provided a Telegram Chinese group retrieval device for fusing multi-source data, including:
[0044] A data collection module, configured to obtain a search term, and perform a Telegram Chinese group search on the search term to generate a multi-source fusion group;
[0045] A feature calculation module, configured to analyze the set of chat records corresponding to the multi-source fusion group, obtain a set of feature words, and screen the multi-source fusion group based on the set of feature words to obtain a conforming feature group V0;
[0046] An associated association module for obtaining an associated group R based on the sharing groups in the corresponding group chat record set that conform to the feature group V t-1 ; performing an associative search on each Telegram Chinese group in the associated group R t-1 to generate an associated association group L t-1 ; screening the associated group R based on the feature words t-1 and the associated association group L t-1 to obtain the feature group V that conforms t-1 ; where t represents the number of iteration rounds; t
[0047] A result generation module for, when the feature group V that conforms is not an empty set, setting t = t + 1 and returning to the associated association module to perform an associative search on each Telegram Chinese group in the associated group R t to generate an associated association group L t-1 ; when the feature group V that conforms is an empty set, obtaining a Telegram Chinese group retrieval result based on the multi-source fusion group and the set V of feature groups that conform; where V = {V0,..., V t-1 t}. t-1
[0048] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, characterized by including:
[0049] A processor;
[0050] A memory for storing executable instructions of the processor;
[0051] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the above-mentioned Telegram Chinese group retrieval method for fusing multi-source data in any one of the above.
[0052] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, characterized in that when the program instructions are executed by a processor, the above-mentioned Telegram Chinese group retrieval method for fusing multi-source data in any one of the above is implemented.
[0053] The method proposed by the present invention has the following advantages and effects:
[0054] 1) It can effectively utilize existing network resources and provide a more comprehensive Chinese group retrieval ability.
[0055] 2) It can effectively retrieve groups that are irrelevant to the title but relevant to the group chat.
[0056] 3) There is no need to conduct extensive searches and frequent updates on groups. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 is the overall flowchart of a Telegram group retrieval that integrates multi-source data provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] To make the technical solutions of the present invention more obvious and understandable, specific embodiments are given and described in detail below in conjunction with the accompanying drawings.
[0059] This embodiment discloses a method and system for Telegram group retrieval that integrates multi-source data, and the specific description is as follows:
[0060] The overall process of the overall architecture of the system is as Figure 1 shown, and this embodiment includes the following steps:
[0061] Step 1: Use multiple data sources to perform multi-source data retrieval on the search term, comprehensively obtain the groups from the multiple data sources, and perform duplicate removal processing on the groups in the set.
[0062] The multiple data source searches described above include: Google search, Twitter targeted data crawler, and group search using a Telegram Chinese retrieval robot.
[0063] Among them, when using Google Custom Search, the retrieval scope of the data is preferably set to telegram.org first, and combined with the search term for retrieval, so as to obtain the qualified results in the Google data source;
[0064] Among them, when using the Twitter targeted data crawler, that is, use the crawler technology to search Twitter for tweets that contain the keyword and the Telegram group link, and capture the Telegram group data. Here, the characteristic field of the Telegram group chat link is: " https: / / t.me / ", and the English string following the characteristic field is the group ID;
[0065] Among them, when using the Telegram Chinese retrieval robot for group retrieval, that is, use the message sending and receiving interface provided by Telegram to ask the Telegram Chinese retrieval robot the keyword to be queried, and use the receiving interface to listen to the results of the Telegram Chinese retrieval robot, and extract the group ID in the results. Here, the Telegram Chinese retrieval robot Q&A is that different developers use their respective technical means to provide a question-and-answer retrieval service for other users through the Telegram robot account with the Chinese retrieval function they have implemented.
[0066] Step 2: Use the Telegram interface to perform an associative search on the search terms.
[0067] The associative search mentioned above means using the English group search interface provided by Telegram to search for the pinyin of the search terms, obtaining a list of associated groups for the search terms, so as to achieve an effective associative extended search for groups with group names, group IDs, and the pinyin structure of the search terms. This group list includes groups where the group username is equal to the search pinyin, as well as a list of groups where the group username is similar to the pinyin of the search terms, or a list of groups where the group title contains the pinyin of the search terms.
[0068] In another example, during the search process, it is not limited to searching for the search terms input by the user, but comprehensively searches in combination with the multi-source data search results that have been obtained.
[0069] Step 3: Merge the multi-source data search results and the associative search results, and remove duplicates from the merged group list to obtain the first-phase multi-source fusion group list.
[0070] Step 4: Analyze the feature words in the chat records of the searched groups.
[0071] The analysis of the feature words in the group chat mentioned above means using the keywords that appear in the group chat and the order of the keywords in the conversation to construct a keyword graph, and screening out the feature words through the keyword link strength.
[0072] The keyword graph mentioned above means extracting keywords from the chat records for each group. The keywords are used as nodes, and edges are made between the keywords that appear simultaneously. The keywords of all groups finally form a unified graph structure to obtain the feature word relationship graph. Specifically, the feature word relationship graph means constructing a keyword association graph using the keywords that appear in the group chat and the order of the keywords in the conversation; among them, the keywords are first obtained by using word segmentation technology to perform word segmentation on the group chat data and using word frequency statistics technology; the construction of the association graph is a keyword relationship graph constructed using the keywords and the order of the keywords in the conversation, reflecting the pointing relationship between words and the strength of the link between keywords.
[0073] In the construction of the graph structure, when the same keyword appears repeatedly, the count of its node is increased, and when the same keyword pair appears repeatedly, the count of the edge between the nodes is increased.
[0074] Finally, select the keywords with higher node counts in the graph structure and the nodes with higher association degrees with them as the feature words.
[0075] Among them, the group chat information is obtained through the iter_messages interface provided in the Telegram open-source package. This interface can be used to traverse and query the content of specific group chats, and the groups to be queried are all the groups included in the multi-source fusion group list.
[0076] Step Five: Filter out irrelevant groups using feature words.
[0077] That is, use the feature words of the groups obtained from the group chat information in Step Four to reverse-check the group chat content of the groups in the multi-source group list in the first stage. According to the matching degree between the group chat keywords and the feature words, filter out the groups whose group chat content does not match the features, and only retain the groups whose group chat content matches the feature words. Here, a multi-source fusion group list after feature word filtering is obtained.
[0078] The aforementioned matching degree is obtained based on the number of feature words included in the group chat keywords. The more feature words are matched, the higher the matching degree.
[0079] Step Six: Conduct associated searches for the groups that meet the features.
[0080] Screen all shared groups from the group chat information and regard them as the associated group list of the groups. Remove the duplicates in the associated list itself, as well as the duplicates between the associated list and the multi-source fusion list.
[0081] The aforementioned associated search means using the group chat shares involved in the group chat records to discover other groups associated with the current group, and all the shared groups form the shared group list of the current group.
[0082] After completing the associated searches for all groups, perform a deduplication operation on the overall associated groups. When deduplicating, it includes the duplicates of the associated groups of each group, as well as the duplicates between the associated groups and the multi-source fusion list.
[0083] Step Seven: Conduct associative searches for the associated results.
[0084] Conduct associative searches for each of the associated groups obtained in Step Six one by one, and incorporate the newly obtained groups after association into the associated group list. After completing the group associative searches, perform a deduplication operation on the updated associated group list.
[0085] In one example, for the discovery of the associated groups, during the associated discovery process, not only retrieve and analyze the associated groups included in the group chat, but also conduct associative expansion searches for each associated group simultaneously. After completing the complete associative expansion search, perform feature word analysis on the obtained groups to filter out the list whose group chat feature words do not meet the feature word relationship diagram; for the newly obtained associated groups, repeat this associated discovery process until no new associated groups that meet the characteristics of the group chat feature words are found.
[0086] Step Eight: Filter out the irrelevant groups in the associated association list obtained in Step Seven using feature words.
[0087] Using the group chat information retrieval interface, obtain the group chat information for the groups in the associated group list. Combine with the feature words obtained in Step Four to filter the groups in the association list, filter out the groups whose group chat content does not conform to the features, and retain the groups whose group chat content conforms to the features to obtain an associated group list that conforms to the features.
[0088] Step Nine: Repeat Steps Six to Eight.
[0089] Continue to perform the associated searches involved in Steps Six and Seven on the obtained associated group list, and filter out the groups whose group chat content conforms to the features by Step Eight. Repeat these three steps continuously to expand the scope of the associated search until the group chat content of the newly emerged groups in the obtained associated groups no longer conforms to the features.
[0090] Step Ten: Merge all the results.
[0091] Merge the multi-source fusion group list obtained in Step Three with the final associated group list to obtain a complete retrieval list.
[0092] The above embodiments are only used to illustrate the technical solutions of the invention and are not intended to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the invention. The protection scope of the present invention shall be subject to what is described in the claims.
Claims
1. A method for retrieving Chinese Telegram groups by integrating multi-source data, characterized in that, The method includes: Obtain a search term, and perform a Telegram Chinese group search on the search term to generate a multi-source fusion group; Analyze the set of group chat records corresponding to the multi-source fusion group to obtain a set of feature words, and filter the multi-source fusion group based on the set of feature words to obtain a conforming feature group V0; Based on the sharing group in the corresponding group chat record set that conforms to the feature group V t-1 obtain the associated group R t-1 ; where t represents the number of iteration rounds; Perform an associative search on each Telegram Chinese group in the associated group R t-1 to generate an associated associative group L t-1 ; Filter the associated group R based on the characteristic words t-1 and the associated association group L t-1 , to obtain the characteristic group V that meets the requirements t ; In the case where the compliance feature group V t is not an empty set, let t = t + 1, and return to the association group R t-1 to perform an associative search on each Telegram Chinese group in it, generating an associated associative group L t-1 ; In the case where the compliance feature group V t is an empty set, a Telegram Chinese group retrieval result is obtained based on the multi-source fusion group and the set V of compliance feature groups; where V = {V0, …, V t-1}; The generation of the multi-source fusion group includes: Use multiple data sources to perform a Telegram Chinese group search on the search term to obtain a multi-source data search group; the multiple data sources include: Google data source, Twitter data source, and other third-party Telegram group information retrieval service data sources. The other third-party Telegram group information retrieval service data source is a data source obtained by using a Telegram Chinese retrieval robot for group search. The obtaining of the multi-source data search group includes: Adopt a custom search mode to directionally search for the search term within the scope of telegram.org to obtain the search results corresponding to the Google data source; Use web crawler technology to directionally search for the search term in Twitter data and filter the data containing the Telegram group field to obtain the search results corresponding to the Twitter data source; Through the question-and-answer service of the Telegram robot account in the other third-party Telegram retrieval service, search for the search term to obtain the search results corresponding to the other third-party Telegram group information retrieval service data source; Merge the search results corresponding to the Google data source, the search results corresponding to the Twitter data source, and the search results corresponding to the other third-party Telegram group information retrieval service data source, and perform deduplication to obtain a multi-source data search group; Based on the English group retrieval interface provided by Telegram, perform a Telegram Chinese group search on the pinyin of the search term and the pinyin similar to the pinyin of the search term to obtain a search term association group; Merge the multi-source data retrieval group and the retrieval term association group, and perform deduplication to obtain a multi-source fusion group; generate the associated association group L t-1 , including: Obtain the associated group R t-1 The group name of each Telegram Chinese group in it; Generate an approximate group name similar to the group name and obtain the pinyin of the approximate group name; Perform Telegram Chinese group search based on multiple data sources, and / or perform Telegram Chinese group search on the pinyin of the approximate group name based on the English group retrieval interface provided by Telegram, so as to obtain the associated and associated group L t-1 。 2. The method according to claim 1, wherein The analysis of the set of group chat records corresponding to the multi-source fusion group to obtain a set of feature words includes: For the multi-source fusion group, use word segmentation technology to segment the group chat records of each Telegram Chinese group, use the high-frequency words in the segmentation results as keywords, and generate keyword pairs based on the order of the keywords in the conversation; Use the keywords in the keyword pairs as nodes and the association of the keywords in the keyword pairs as edges to construct a keyword relationship graph; among them, the weight of the node is the number of times the keyword appears, and the weight of the edge is the number of times the keyword pair appears; Filter the keywords based on the weight of the nodes to obtain a set of main feature words; According to the main feature words and the weight of the edges connecting the main feature words, obtain a set of auxiliary feature words; Merge the set of main feature words and the set of auxiliary feature words to obtain a set of feature words.
3. A Telegram Chinese group retrieval device for integrating multi-source data based on the method described in claim 1 or 2, characterized in that, The device includes: A data collection module, configured to obtain a search term and perform a Telegram Chinese group search on the search term to generate a multi-source fusion group; A feature calculation module, configured to analyze a set of group chat records corresponding to the multi-source fusion group, obtain a set of feature words, and filter the multi-source fusion group based on the set of feature words to obtain a conforming feature group V0; An associated association module for obtaining an associated group R based on the sharing groups in the corresponding group chat record set that conform to the feature group V t-1 ; performing an associative search on each Telegram Chinese group in the associated group R to generate an associated association group L t-1 ; filtering the associated group R t-1 with the associated association group L t-1 to obtain the feature group V that meets the requirements t-1 ; where t represents the number of iteration rounds t-1 ; t A result generation module, which is used to, when the compliant feature group V t is not an empty set, set t = t + 1, and return to the associated association module to perform an associative search on each Telegram Chinese group in the associated group R t-1 to generate an associated association group L t-1 ; when the compliant feature group V t is an empty set, obtain a Telegram Chinese group retrieval result based on the multi-source fusion group and the set V of compliant feature groups; where V = {V0, …, V t-1}.
4. An electronic device, characterized in that, It includes: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the Telegram Chinese group retrieval method for fusing multi-source data according to any one of claims 1-2.
5. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the Telegram Chinese group retrieval method for fusing multi-source data according to any one of claims 1-2 is implemented.
Citation Information
Patent Citations
A method and a device for retrieving information
CN109902152A
Chinese and English paper data classification and query method
CN112632282A