Telegram-oriented dialogue data automatic acquisition and fine-grained personalized information labeling method

Through the automatic collection of dialogue data for Telegram and fine-grained personalized information labeling methods, the problems of inefficient identification of illegal users and high manual labeling in the prior art are solved, efficient and accurate data collection and labeling are achieved, and personalized dialogue models for identifying illegal users are supported.

CN120123433APending Publication Date: 2025-06-10NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510195311.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively identify illegal users on the Telegram platform. The traditional detection methods are inefficient and have problems with high false alarm rates and missed alarm rates. At the same time, the cost of manually labeling illegal users is extremely high, making it difficult to carry out on a large scale.

Method used

Automatic collection of dialogue data for Telegram and fine-grained personalized information annotation method, group dialogue data crawling is used through Telegram client disguise and dynamic proxy pool optimization, and advertising, duplicate information filtering and effective dialogue screening are combined with language models and rule scripts, and fine-grained personalized information annotation is used using small sample prompt technology and large language models.

Benefits of technology

It realizes efficient data collection and labeling, reduces the cost of manual labeling, improves the quality and accuracy of conversation data, and supports training of personalized conversation models that can identify illegal users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123433A_ABST
    Figure CN120123433A_ABST
Patent Text Reader

Abstract

The invention discloses a Telegram-oriented dialogue data automatic acquisition and fine-grained personalized information labeling method, and relates to the technical field of data acquisition and labeling, and the method comprises the steps of Telegram dialogue data automatic acquisition, high-quality dialogue data extraction and fine-grained personalized information labeling, and the acquisition of Telegram dialogue data is realized through a client camouflage technology. The distributed crawler can bypass access control and forbidding limitation of Telegram, so that historical dialogue data of the target group can be continuously and stably obtained; in the high-quality dialogue data extraction, effective dialogues are screened through predefined rule scripts and training models, and invalid information and redundant contents are filtered out; the fine-grained personalized information labeling is divided into two stages. According to the method, efficient data collection and accurate data extraction are achieved, a few-sample prompting technology and a large language model are utilized, manual annotation is reduced, a feedback iteration mechanism is optimized, annotation accuracy and consistency are improved, and high-quality personalized dialogue data are generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data collection and annotation, and more specifically to a method for automatically collecting dialogue data and fine-grained personality information annotation for Telegram in the technical field. Background Art

[0002] Telegram is one of the most popular instant messaging tools at present. Due to its emphasis on security and privacy and high real-time characteristics, it has attracted a large number of users to download and use. However, its privacy protection features have also become a cover for some illegal activities, posing an increasingly severe threat to social security. These illegal users often disguise themselves as ordinary users and hide in a vast amount of dialogue data, making it difficult to effectively identify them through traditional detection means. Existing illegal user detection methods mainly rely on manual reporting, keyword monitoring, or behavior analysis. These methods are not only inefficient but also tend to have a high false positive rate and false negative rate in the face of complex dialogue scenarios. At the same time, the concealment of illegal users and the huge amount of data on social media platforms make the cost of manually annotating illegal users extremely high and difficult to carry out on a large scale. In addition, due to Telegram's encryption protection of communications, external monitoring and analysis are more difficult.

[0003] To effectively address this challenge, the present invention begins to focus on using the generative dialogue technology of large language models to identify illegal users by actively participating in conversations. However, the effectiveness of the generative dialogue model depends on high-quality annotated data, especially on the fine-grained annotation of personality dialogue data in this scenario. To detect illegal users, the system needs to train personality dialogue models that can disguise as different users. These models can not only have natural conversations with illegal users but also obtain key information from the conversations, ultimately achieving the identification and monitoring of illegal users. Traditional data annotation methods are inefficient when annotating a vast amount of dialogue data and are difficult to capture fine-grained personality characteristics when dealing with personality conversations. Existing automated annotation technologies often require a large amount of high-quality training data when generating personality annotations and perform poorly in the extraction of personality information and the understanding of dialogue context. Therefore, there is an urgent need for an efficient data annotation method that can guide the generation of large-scale automated annotations through a small number of manual annotation examples and ultimately provide support for training personality dialogue models. Summary of the Invention

[0004] The purpose of the present invention is: to solve the above technical problems, the present invention provides a method for automatically collecting dialogue data and fine-grained personality information annotation for Telegram.

[0005] The present invention specifically adopts the following technical solutions to achieve the above purpose:

[0006] The present invention provides a method for automatically collecting Telegram conversation data and fine-grained personality information annotation, including the following steps:

[0007] S1. Automatic collection of Telegram conversation data, including the following methods:

[0008] S11. Disguise of Telegram client;

[0009] S12. Crawling of group conversation data optimized by a dynamic proxy pool;

[0010] S2. Extraction of high-quality conversation data, including the following methods:

[0011] S21. Filtering of advertising messages;

[0012] S22. Filtering of duplicate messages;

[0013] S23. Screening of valid conversations;

[0014] S3. Fine-grained personality information annotation: Fine-grained personality information annotation is used to extract and annotate multi-dimensional features in the conversations of target users, ensuring that the trained personality conversation model accurately reflects the personality characteristics of users: The annotation process is divided into three stages:

[0015] S31. Manual annotation stage;

[0016] S32. Automatic annotation stage assisted by large language models;

[0017] S33. Optimization stage combined with a feedback mechanism to improve the accuracy and consistency of annotation;

[0018] S4. Generation of annotation results: After the fine-grained personality information annotation process in step S3 is completed, a complete personality conversation dataset is generated.

[0019] In one implementation, in step S11, the specific steps of disguising the Telegram client are as follows:

[0020] S111. Analyze the source code of the Telegram desktop client to find the key fields that affect client property authentication, including api_id, api_hash, and lang_pack. These three fields play a role in identifying the client during the authentication process; replace these fields during authentication to make the behavior of the crawler module consistent with that of the official client, thereby avoiding the monitoring and banning of Telegram for non-official clients;

[0021] After applying for the api_id and api_hash, automatically replace the parameters during each login to avoid being recognized as a non-official client. By analyzing the request messages in the MTProto protocol, ensure that the api_id, api_hash, and lang_pack are consistent with the official client during the authentication process;

[0022] S113. Use Python libraries such as Telethon to interact with the Telegram server, and dynamically replace relevant fields during the authentication process to simulate the behavior of the official client and ensure that the abnormal detection mechanism is not triggered during the data collection process.

[0023] In one implementation, in step S12, the specific steps for optimizing the crawling of group conversation data by the dynamic proxy pool are as follows:

[0024] S121. As Figure 2 , when crawling a large amount of conversation data, to avoid IP bans and improve access efficiency, use a dynamic proxy pool;

[0025] S122. During the crawling process, the URLs, conversation histories, and user information of the groups are first cached in Redis for temporary storage and load balancing, and then regularly transferred to the MySQL database for subsequent processing and analysis. The optimization strategy of the dynamic proxy pool is based on the following formula:

[0026]

[0027] In the formula, T proxy represents the usage interval time of the proxy, P i is the historical usage success rate of the proxy, and R i is the remaining usage times of the current proxy; this formula dynamically adjusts the rotation frequency of the proxy, maximizes the data collection efficiency while avoiding triggering the IP ban mechanism of Telegram, and selects the proxy IP with the highest usage success rate and the most remaining times for each request, dynamically adjusting the access strategy to balance the collection efficiency and the ban risk.

[0028] In one implementation, in step S21, the advertisement message filtering package completes the filtering of advertisement messages through a pre-trained language-independent sentence representation model and a neural network classification model. This process uses cross-language text encoding and feature extraction techniques to identify and filter advertisement messages from the conversation data. Specifically, the feature extraction steps include the following:

[0029] S211. Extraction of the number of "@" symbols: Advertisement messages usually contain multiple "@" symbols to attract attention. Count the number of "@" symbols in the message and use it as one of the features;

[0030] S212. Extraction of the number of links: Advertising messages usually contain multiple hyperlinks. Count the number of URLs in the message. Frequently appearing hyperlinks are common features of advertisements.

[0031] S213. Extraction of the number of Emoji symbols: Advertising messages often use excessive Emoji symbols to enhance expressiveness. Identify and count the number of Emoji symbols in the message and use it as a classification feature.

[0032] S214. Extraction of the message length: Advertising messages are usually longer than ordinary messages. Extract the message length by calculating the number of characters in the message to help distinguish advertisements from ordinary messages.

[0033] S215. Extraction of the number of line breaks: Advertising messages often separate content with multiple line breaks. Count the number of line breaks as a basis for identifying advertising messages.

[0034] S216. After cross - linguistically encoding the message using a pre - trained language - independent sentence representation model, use a feed - forward neural network to classify the extracted features to determine whether the message is an advertisement.

[0035] In one implementation, in step S22, duplicate message filtering includes the following:

[0036] S221. Identify and filter duplicate or highly similar messages based on the bag - of - words model and cosine similarity algorithm. Each message is vectorized using the bag - of - words model and a vector is generated based on the frequency of word occurrences. The similarity calculation formula is as follows:

[0037]

[0038] In the formula, X and Y are text vectors, n is the length of the text vector, X i and Y i are the values of the i - th position in vectors X and Y respectively;

[0039] S222. If the similarity of two messages exceeds the set threshold, it is determined as a duplicate message and filtered out.

[0040] In one implementation, in step S23, valid conversation screening: After filtering advertising messages and cleaning duplicate messages, screen for valid conversations with depth and user personality characteristics through a rule script. The specific steps are as follows:

[0041] S231. Screen non - robot users: Analyze the identities of conversation participants and screen out conversations participated by at least two non - robot users. Determine and exclude conversations dominated by robots through user behavior patterns or client identifiers.

[0042] S232. Dialogue Turn Detection: Count the interaction turns of each dialogue and filter out the dialogues with more than 10 interaction turns. Detect message back-and-forth through timestamps and user IDs, and filter out the dialogues with less than 10 interaction turns;

[0043] S233. User Activity Judgment: Analyze the activity of dialogue participants and filter out the users who have participated in at least 5000 dialogues. Record the dialogue history of users and count the total number of dialogues, and exclude the users who do not meet the activity requirements.

[0044] In one implementation, in step S31, the manual annotation stage includes the following steps:

[0045] In the initial stage, randomly select some dialogue data for manual analysis and annotation. The annotation content is related to the target user. Filter out the representative dialogue segments that are closely related to the user's personality, and extract the multi-dimensional personality information of the user. The specific annotation content includes but is not limited to the user's name, gender, identity background, behavior characteristics, language style, personality views, etc. The information in each dimension will be detailedly annotated to ensure that the annotated data has sufficient granularity and comprehensiveness. Finally, generate detailed personality information segments, and these annotated data are saved as high-quality annotated data for guiding the next automated annotation.

[0046] In one implementation, in step S32, the automated annotation stage includes the following steps:

[0047] S321. Based on a small amount of high-quality annotated data generated by manual annotation, the system inputs this data into a large language model through few-shot prompting techniques to guide the automated annotation of large-scale unannotated data. Through few-shot prompting techniques, the model can learn and extract the personality characteristics of users, extract multi-dimensional information in the user's dialogue during the automated annotation process, and automatically remove topics irrelevant to the annotation. At this stage, the model will batch process large-scale unannotated data to generate automated annotation data. These automated annotation data include detailed information related to the user's personality, similar to the multi-dimensional personality information segments of manual annotation. After the automated annotation is completed, the system will calculate a confidence score for each annotation result using the following formula:

[0048]

[0049] In the formula, z is the output value predicted by the model, which is converted into a confidence score between 0 and 1 through the Sigmoid function. The system can screen the annotation results according to the confidence and retain the data with high confidence.

[0050] S322. Randomly select a part of the data generated from the automatic annotation for review by experts. The experts manually check these annotation results and provide feedback based on incorrect or inaccurate annotations. The feedback results will directly affect the annotation prompts of the large language model and adjust the annotation parameters. Iterate according to the feedback. Through multiple optimizations of the annotation prompts and parameter adjustments, gradually improve the overall accuracy of the automatic annotation. As the iteration progresses, the similarity between the automatic annotation results and the manual annotation results gradually increases, and finally ensure the consistency between the machine-annotated data and the manually-annotated data.

[0051] In one embodiment, in step S4, the generation of the annotation results includes the following parts:

[0052] S41. User-level personalized dialogue data: including personality information, dialogue scenarios, and dialogue responses, comprehensively capturing the fine-grained personality characteristics of users in specific scenarios;

[0053] S42. General personalized dialogue data: composed of personality information fragments, dialogue scenarios, and dialogue responses, reflecting the common personality characteristics across users;

[0054] S43. Personalized sparse dialogue data: including personality information fragments, dialogue scenarios, and personalized sparse dialogue responses, used to enhance the model's understanding of personality information fragments;

[0055] S44. Personalized inconsistent dialogue data: generated by combining dialogue responses with mismatched personality information fragments, used to train the model to identify dialogue situations with inconsistent personalities.

[0056] The beneficial effects of the present invention are as follows:

[0057] 1. Efficient data collection is achieved by disguising through the Telegram client and scheduling with a dynamic proxy pool, effectively bypassing access restrictions and enabling large-scale, continuous and stable data collection.

[0058] 2. Precise data extraction combines a language model and rule scripts to automatically filter out noise data such as advertisements and duplicate information, ensuring that the extracted dialogue data is efficient and relevant.

[0059] 3. Fine-grained personality annotation uses few-shot prompting techniques and large language models to reduce manual annotation, optimize the feedback iteration mechanism, improve annotation accuracy and consistency, and generate high-quality personalized dialogue data. Brief Description of the Drawings

[0060] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.

[0061] Figure 1 is the overall flowchart of Telegram automated data collection and personality annotation of the present invention.

[0062] Figure 2 is the flowchart of automatic collection of Telegram conversation data.

[0063] Figure 3 is the composition diagram of the annotated Telegram personality conversation dataset. Detailed implementation manners

[0064] To make the technical problems, technical solutions and technical effects of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0065] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0066] Embodiment 1

[0067] As Figure 1 shown, this embodiment provides a method for automatic collection of Telegram conversation data and fine-grained personality information annotation, including the following steps:

[0068] S1. Automatic collection of Telegram conversation data, including the following two steps of crawling group conversation data with Telegram client disguise and dynamic proxy pool optimization:

[0069] The specific steps of S11. Telegram client disguise are as follows:

[0070] S111. Analyze the source code of the Telegram desktop client to find the key fields that affect client property authentication, including api_id, api_hash, and lang_pack. These three fields play a role in identifying the client during the Telegram authentication process. Replace these fields during authentication to make the behavior of the crawler module consistent with that of the official client, thereby avoiding Telegram's monitoring and banning of unofficial clients.

[0071] S112. After applying for api_id and api_hash, automatically replace the parameters during each login to avoid being recognized as an unofficial client. By analyzing the request messages in the MTProto protocol, ensure that api_id, api_hash, and lang_pack are consistent with the official client during the authentication process.

[0072] S113. Use Python libraries such as Telethon to interact with the Telegram server and dynamically replace relevant fields during the authentication process to simulate the behavior of the official client and ensure that the anomaly detection mechanism is not triggered during the data collection process.

[0073] S12. The specific steps for optimizing the crawling of group chat data using a dynamic proxy pool are as follows:

[0074] S121. As Figure 2 , when crawling a large amount of chat data, to avoid IP bans and improve access efficiency, use a dynamic proxy pool.

[0075] S122. During the crawling process, the URLs of the groups, chat histories, and user information are first cached in Redis for temporary storage and load balancing, and then regularly transferred to the MySQL database for subsequent processing and analysis. The optimization strategy of the dynamic proxy pool is based on the following formula:

[0076]

[0077] In the formula, T proxy represents the usage interval time of the proxy, P i is the historical usage success rate of the proxy, and R i is the remaining usage times of the current proxy. This formula dynamically adjusts the rotation frequency of the proxy, maximizing the data collection efficiency while avoiding triggering Telegram's IP ban mechanism. For each request, select the proxy IP with the highest usage success rate and the most remaining times, and dynamically adjust the access strategy to balance the collection efficiency and the ban risk.

[0078] The specific algorithm flow is as follows:

[0079]

[0080]

[0081] S2. High-quality dialogue data extraction, including three steps: advertisement message filtering, duplicate message filtering, and effective dialogue screening:

[0082] S21. Advertisement message filtering package: Advertisement messages are filtered through a pre-trained language-independent sentence representation model and a neural network classification model. This process utilizes cross-lingual text encoding and feature extraction techniques to identify and filter advertisement messages from dialogue data. Specifically, the feature extraction step includes the following:

[0083] S211. Extraction of the number of "@" symbols: Advertisement messages usually contain multiple "@" symbols to draw attention. Count the number of "@" symbols in the message and use it as one of the features.

[0084] S212. Extraction of the number of links: Advertisement messages usually contain multiple hyperlinks. Count the number of URLs in the message. Frequently occurring hyperlinks are common features of advertisements.

[0085] S213. Extraction of the number of Emoji symbols: Advertisement messages often use excessive Emoji symbols to enhance expressiveness. Identify and count the number of Emoji symbols in the message and use it as a classification feature.

[0086] S214. Extraction of the message length: Advertisement messages are usually longer than ordinary messages. Extract the message length by calculating the number of characters in the message to help distinguish advertisements from ordinary messages.

[0087] S215. Extraction of the number of line breaks: Advertisement messages often separate content with multiple line breaks. Count the number of line breaks as a basis for identifying advertisement messages.

[0088] S216. After cross-lingually encoding the message using a pre-trained language-independent sentence representation model, use a feed-forward neural network to classify the extracted features to determine whether the message is an advertisement.

[0089] S22. Duplicate message filtering includes the following:

[0090] S221. Identify and filter duplicate or highly similar messages based on the bag-of-words model and the cosine similarity algorithm. Each message is vectorized through the bag-of-words model and a vector is generated based on the frequency of word occurrences. The similarity calculation formula is as follows:

[0091]

[0092] In the formula, X and Y are text vectors, n is the length of the text vector, X i and Y iThey are the values of the i-th bit in vectors X and Y respectively;

[0093] S222. If the similarity of two messages exceeds the set threshold, they are determined to be duplicate messages and filtered.

[0094] The specific process of the filtering algorithm is as follows:

[0095]

[0096]

[0097] S23. Effective dialogue screening: After filtering advertisement messages and cleaning duplicate messages, effective dialogues with depth and user personality characteristics are screened through rule scripts. The specific steps are as follows:

[0098] S231. Screening non-robot users: Analyze the identities of dialogue participants, screen out dialogues participated by at least two non-robot users, and determine and eliminate robot-dominated dialogues through user behavior patterns or client identifiers;

[0099] S232. Dialogue turn detection: Count the interaction turns of each dialogue, and screen out dialogues with more than 10 interaction turns. Detect message back-and-forth through timestamps and user IDs, and dialogues with less than 10 interaction turns will be filtered;

[0100] S233. User activity judgment: Analyze the activity of dialogue participants, and screen out users who have participated in at least 5000 dialogues. Record the dialogue history of users and count the total number of dialogues, and exclude users who do not meet the activity requirements.

[0101] S3. Fine-grained personality information annotation: Fine-grained personality information annotation is used to extract and annotate multi-dimensional features in the target user's dialogue to ensure that the trained personality dialogue model accurately reflects the user's personality characteristics. The annotation process is divided into three stages:

[0102] S31. The manual annotation stage includes the following steps:

[0103] In the initial stage, randomly select some dialogue data for manual analysis and annotation. The annotation content is related to the target user. Screen out representative dialogue fragments that are closely related to the user's personality, and extract the user's multi-dimensional personality information. The specific annotation content includes but is not limited to the user's name, gender, identity background, behavior characteristics, language style, personality views, etc. Information in each dimension will be annotated in detail to ensure that the annotated data has sufficient granularity and comprehensiveness. Finally, generate detailed personality information fragments. These annotated data are saved as high-quality annotated data for guiding the next automated annotation.

[0104] In S32, the automated annotation stage includes the following steps:

[0105] S321. Based on a small amount of high-quality labeled data generated by manual annotation, the system inputs this data into the large language model through few-shot prompting technology to guide the automatic annotation of large-scale unlabeled data. Through few-shot prompting technology, the model can learn and extract the user's personality characteristics, extract multi-dimensional information in the user's conversation during the automatic annotation process, and automatically remove topics irrelevant to the annotation. At this stage, the model will batch process large-scale unlabeled data to generate automatically labeled data, which includes detailed information related to the user's personality, similar to multi-dimensional personality information segments of manual annotation. After the automatic annotation is completed, the system will calculate a confidence score for each annotation result using the following formula:

[0106]

[0107] In the formula, z is the output value predicted by the model, which is converted into a confidence score between 0 and 1 through the Sigmoid function. The system can screen the annotation results according to the confidence and retain the data with high confidence.

[0108] S322. Randomly select a part of the data from the dataset generated by automatic annotation for review by experts. The experts manually check these annotation results and provide feedback based on incorrect or inaccurate annotations. The feedback results will directly affect the annotation prompt content of the large language model and adjust the annotation parameters. Iterate according to the feedback. Through multiple optimizations of the annotation prompt and parameter adjustment, gradually improve the overall accuracy of automatic annotation. As the iteration progresses, the similarity between the automatic annotation results and the manual annotation results gradually increases, and finally ensure the consistency between the machine-labeled data and the manual-labeled data.

[0109] The specific process of the fine-grained personality information annotation algorithm is as follows:

[0110]

[0111]

[0112]

[0113] S4. Generation of annotation results: After the fine-grained personality information annotation process in step S3 is completed, a complete personality conversation dataset is generated, such as Figure 3 , which includes the following parts:

[0114] S41. User-level personality conversation data: including personality information, conversation scenarios, and conversation responses, comprehensively capturing the user's fine-grained personality characteristics in specific scenarios;

[0115] S42, General Personality Dialogue Data: Composed of personality information fragments, dialogue scenarios, and dialogue responses, reflecting the common personality characteristics across users;

[0116] S43, Personality Sparse Dialogue Data: Including personality information fragments, dialogue scenarios, and personality sparse dialogue responses, used to enhance the model's understanding of personality information fragments;

[0117] S44, Personality Inconsistent Dialogue Data: Generated by combining dialogue responses with mismatched personality information fragments, used to train the model to recognize dialogue situations with personality inconsistencies.

Claims

1. A method for automatic collection of conversation data and fine-grained personal information annotation for Telegram, characterized in that: The steps include: S1. Automatic collection of Telegram conversation data, including the following methods: S11, Telegram client disguise; S12, group conversation data crawling with dynamic proxy pool optimization; S2. Extraction of high-quality conversation data, including the following methods: S21, advertising message filtering; S22, duplicate message filtering; S23, effective dialogue screening; S3. Fine-grained personality information annotation: Fine-grained personality information annotation is used to extract and annotate multi-dimensional features in the target user's conversation to ensure that the trained personality conversation model accurately reflects the user's personality characteristics: The annotation process is divided into three stages: S31, manual labeling stage; S32, automatic annotation stage assisted by large language model; S33, in order to improve the accuracy and consistency of annotation, an optimization phase is carried out in combination with a feedback mechanism; S4. Generation of annotation results: After the fine-grained personality information annotation process in step S3 is completed, a complete personality dialogue dataset is generated.

2. According to claim 1, a method for automatic collection of conversation data and fine-grained personality information annotation for Telegram, characterized in that: In step S11, the specific steps of Telegram client disguise are as follows: S111. Analyze the source code of the Telegram desktop client and find the key fields that affect the client attribute authentication, including api_id, api_hash and lang_pack. These three fields play the role of identifying the client in the authentication process of Telegram. Replace these fields during authentication to make the crawler module behavior consistent with the official client, thereby circumventing Telegram's monitoring and banning of unofficial clients. S112. After applying for api_id and api_hash, automatically replace the parameters at each login to avoid being identified as an unofficial client. By analyzing the request message in the MTProto protocol, ensure that api_id, api_hash and lang_pack are consistent with the official client during the authentication process; S113. Use the Python library to interact with the Telegram server and dynamically replace relevant fields during the authentication process to simulate the official client behavior and ensure that the anomaly detection mechanism is not triggered during data collection.

3. According to claim 1, a method for automatic collection of conversation data and fine-grained personality information annotation for Telegram, characterized in that: In step S12, the specific steps of crawling group conversation data for dynamic proxy pool optimization are as follows: S121. When crawling conversation data on a large scale, a dynamic proxy pool is used to avoid IP blocking and improve access efficiency; S122. During the crawling process, the group URL, conversation history and user information are first cached in Redis for temporary storage and load balancing, and then regularly transferred to the MySQL database for subsequent processing and analysis. The optimization strategy of the dynamic proxy pool is based on the following formula: Where, T proxy represents the usage interval of the agent, P i is the historical success rate of the agent, R i is the remaining number of uses of the current proxy; this formula dynamically adjusts the rotation frequency of the proxy to maximize data collection efficiency while avoiding triggering Telegram's IP ban mechanism. Each time a request is made, the proxy IP with the highest success rate and the most remaining times is selected, and the access strategy is dynamically adjusted to balance collection efficiency and ban risk.

4. According to claim 1, a method for automatic collection of conversation data and fine-grained personality information annotation for Telegram, characterized in that: In step S21, the advertisement message filtering package is filtered by using a pre-trained language-independent sentence representation model and a neural network classification model. This process uses cross-language text encoding and feature extraction technology to identify and filter advertisement messages from conversation data. Specifically, the feature extraction step includes the following steps: S211, extracting the number of "@" symbols: an advertisement message usually contains multiple "@" symbols to attract attention, and the number of "@" symbols in the message is counted and used as one of the features; S212, extracting the number of links: the advertisement message usually contains multiple hyperlinks, and the number of URLs in the message is counted. Frequently appearing hyperlinks are a common feature of advertisements; S213, extraction of the number of Emoji symbols: advertising messages often use too many Emoji symbols to enhance the expressiveness. The number of Emoji symbols in the message is identified and counted, and used as a classification feature; S214, extracting the message length: advertising messages are usually longer than ordinary messages. The message length is extracted by calculating the number of characters in the message to help distinguish advertising from ordinary messages. S215, extracting the number of line breaks: advertising messages often use multiple line breaks to separate content, and the number of line breaks is counted as a basis for identifying the advertising message; S216. After cross-language encoding of the message using a pre-trained language-independent sentence representation model, the extracted features are classified using a feedforward neural network to determine whether the message is an advertisement.

5. According to claim 1, a method for automatic collection of conversation data and fine-grained personality information annotation for Telegram, characterized in that: In step S22, duplicate message filtering includes the following contents: S221. Based on the bag-of-words model and the cosine similarity algorithm, duplicate or highly similar messages are identified and filtered. Each message is vectorized using the bag-of-words model, and a vector is generated based on the frequency of word occurrence. The similarity calculation formula is as follows: In the formula, X and Y are text vectors, n is the length of the text vector, and X i With Y i are the values ​​of the i-th position in vectors X and Y respectively; S222: If the similarity between the two messages exceeds a set threshold, they are determined to be duplicate messages and are filtered.

6. According to claim 1, a method for automatic collection of conversation data and fine-grained personality information annotation for Telegram, characterized in that: In step S23, effective dialogue screening: after filtering advertising messages and cleaning duplicate messages, effective dialogues with depth and user personality characteristics are screened through rule scripts. The specific steps are as follows: S231. Screening non-robot users: Analyze the identities of the conversation participants, screen out conversations involving at least two non-robot users, and determine and eliminate robot-led conversations based on user behavior patterns or client identifiers; S232, dialogue round detection: count the interactive rounds of each dialogue, filter out dialogues with more than 10 interactive rounds, detect message round trips by timestamp and user ID, and filter out dialogues with less than 10 interactive rounds; S233, user activity determination: Analyze the activity of conversation participants, screen out users who have participated in at least 5,000 conversations, record the user's conversation history and count the total number of conversations, and exclude users who do not meet the activity requirements.

7. According to claim 1, a method for automatic collection of conversation data and fine-grained personality information annotation for Telegram, characterized in that: In step S31, the manual labeling stage includes the following steps: In the initial stage, some conversation data are randomly selected for manual analysis and labeling. The labeled content is related to the target user. Representative conversation fragments that are closely related to the user's personality are screened out, and the specific labeling content of the user's multi-dimensional personality information is extracted. The information of each dimension will be labeled in detail to ensure that the labeled data is sufficiently granular and comprehensive. Finally, detailed personality information fragments are generated. These labeled data are saved as high-quality labeled data to guide the next step of automated labeling.

8. According to claim 7, a method for automatic collection of conversation data and fine-grained personality information annotation for Telegram, characterized in that: In step S32, the automatic marking stage includes the following steps: S321. Based on a small amount of high-quality annotated data generated by manual annotation, the system inputs this data into the large language model through the few-sample prompting technology to guide the automatic annotation of large-scale unlabeled data. Through the few-sample prompting technology, the model can learn and extract the user's personality characteristics, extract multi-dimensional information in the user's conversation during the automatic annotation process, and automatically remove topics that are not related to the annotation. At this stage, the model will batch process large-scale unlabeled data to generate automatic annotation data. These automatic annotation data include detailed information related to the user's personality, similar to the multi-dimensional personality information fragments annotated by manual annotation. After the automatic annotation is completed, the system will calculate the confidence score for each annotation result, using the following formula for calculation: In the formula, z is the output value predicted by the model, which is converted into a confidence score between 0 and 1 through the Sigmoid function. The system can filter the annotation results according to the confidence and retain the data with high confidence. S322. Randomly select some data from the dataset generated by automatic annotation for review by experts. The experts manually check the annotation results and provide feedback based on incorrect or inaccurate annotations. The feedback results will directly affect the annotation prompt content of the large language model and adjust the annotation parameters. Iterate according to the feedback, and gradually improve the overall accuracy of automatic annotation by optimizing the annotation prompts and adjusting the parameters multiple times. As the iteration proceeds, the similarity between the automatic annotation results and the manual annotation results gradually increases, ultimately ensuring that the machine-annotated data and the manual-annotated data are consistent.

9. The method for automatic collection of conversation data and fine-grained personality information annotation for Telegram according to claim 8, characterized in that: In step S4, the generation of the annotation results includes the following parts: S41, user-level personalized conversation data: including personality information, conversation scenarios and conversation responses, comprehensively capturing the fine-grained personality characteristics of users in specific scenarios; S42, general personality dialogue data: consists of personality information fragments, dialogue scenes and dialogue responses, reflecting the common personality characteristics of different users; S43, sparse personality dialogue data: including personality information fragments, dialogue scenes and sparse personality dialogue responses, used to enhance the model's understanding of personality information fragments; S44, inconsistent personality conversation data: generated by combining conversation responses with mismatched personality information fragments, used to train the model to identify inconsistent personality conversation situations.