Business processing method, related device, equipment and storage medium

By using a text encoding model generated based on comparison learning, the text to be tested and abnormal text is encoded, and the problem of low detection accuracy in the prior art is solved, and higher text detection accuracy and account detection reliability are achieved.

CN120020856APending Publication Date: 2025-05-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311551408.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-20

Smart Images

  • Figure CN120020856A_ABST
    Figure CN120020856A_ABST
Patent Text Reader

Abstract

The invention discloses a business processing method, a related device, equipment and a storage medium, and particularly relates to natural language processing and machine learning. The method comprises the steps of obtaining a to-be-tested text set; coding each to-be-tested text in the to-be-tested text set by adopting a text coding model to obtain K to-be-tested text vectors; obtaining an abnormal text set; coding each abnormal text in the abnormal text set by adopting a text coding model to obtain T abnormal text vectors; and generating a detection result for the to-be-detected account according to the K to-be-detected text vectors and the T abnormal text vectors. According to the invention, the text coding model generated based on comparative learning is used to obtain the coding vector with better quality, so that the accuracy of text detection can be improved, and the accuracy and reliability of account detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method for business processing, related devices, equipment, and storage media. Background Art

[0002] With the development of Internet technology, more and more users obtain network information by registering accounts on application clients. However, in order to seek huge profits, illegal users often maliciously register some accounts to spread illegal information flows, which greatly affects the user experience. Therefore, how to effectively and accurately detect such illegal accounts is particularly important.

[0003] Currently, the commonly used solution in the industry is as follows: First, obtain the text sent by the account. Then, perform regular matching on the text, that is, find whether there is text with a specific pattern in the text. For example, it can be determined whether malicious words appear in the text. If so, the account that hits the malicious words is blocked.

[0004] However, the inventor found that there are at least the following problems in the current solution: Malicious users can discover which keywords are used as malicious words to match the text through online strategy testing. Therefore, malicious users may use other words to bypass the detection. This results in a low accuracy of text detection, and further leads to inaccurate detection results. Summary of the Invention

[0005] Embodiments of the present application provide a method for business processing, related devices, equipment, and storage media. Using a text encoding model generated based on contrastive learning can obtain better-quality encoding vectors, which is beneficial to improving the accuracy of text detection, and further improving the accuracy and reliability of account detection.

[0006] In view of this, one aspect of the present application provides a method for business processing, including:

[0007] Obtain a set of texts to be tested, where the set of texts to be tested includes K texts to be tested sent by the account to be detected, and K is an integer greater than or equal to 1;

[0008] Encode each text to be tested in the set of texts to be tested using a text encoding model to obtain K text vectors to be tested, where the K text vectors to be tested have a one-to-one correspondence with the K texts to be tested, and the text encoding model is obtained by performing contrastive learning on a positive example encoding vector pair and a negative example encoding vector pair;

[0009] Obtain a set of abnormal texts, where the set of abnormal texts includes T abnormal texts that have been marked as abnormal content, and T is an integer greater than or equal to 1;

[0010] Encode each abnormal text in the abnormal text set using a text encoding model to obtain T abnormal text vectors, where the T abnormal text vectors have a one-to-one correspondence with the T abnormal texts;

[0011] Generate a detection result for the account to be detected based on the K text vectors to be measured and the T abnormal text vectors, where the detection result is used to represent the abnormal degree of the account to be detected.

[0012] On the other hand, this application provides a service processing device, including:

[0013] An acquisition module, configured to acquire a text set to be measured, where the text set to be measured includes K texts to be measured sent by the account to be detected, and K is an integer greater than or equal to 1;

[0014] An encoding module, configured to encode each text to be measured in the text set to be measured using a text encoding model to obtain K text vectors to be measured, where the K text vectors to be measured have a one-to-one correspondence with the K texts to be measured, and the text encoding model is obtained by performing contrastive learning on positive example encoding vector pairs and negative example encoding vector pairs;

[0015] The acquisition module is further configured to acquire an abnormal text set, where the abnormal text set includes T abnormal texts that have been marked as abnormal content, and T is an integer greater than or equal to 1;

[0016] The encoding module is further configured to encode each abnormal text in the abnormal text set using a text encoding model to obtain T abnormal text vectors, where the T abnormal text vectors have a one-to-one correspondence with the T abnormal texts;

[0017] A generation module, configured to generate a detection result for the account to be detected based on the K text vectors to be measured and the T abnormal text vectors, where the detection result is used to represent the abnormal degree of the account to be detected.

[0018] In a possible design, in another implementation manner of the other aspect of the embodiments of this application, the service processing device further includes a processing module and a training module;

[0019] The acquisition module is further configured to acquire a text sample set, where the text sample set includes N text samples, and N is an integer greater than 1;

[0020] A processing module, configured to perform data augmentation processing on each text sample in the text sample set to obtain a positive example encoding vector corresponding to each text sample;

[0021] The acquisition module is further configured to acquire N negative example encoding vectors corresponding to each text sample in the text sample set, where the N negative example encoding vectors are generated based on the N text samples;

[0022] The processing module is further configured to calculate, for each text sample, a loss value corresponding to the text sample according to the positive example encoding vector and N negative example encoding vectors corresponding to the text sample;

[0023] The training module is configured to update the model parameters of the text encoding model to be trained according to the loss value corresponding to each text sample in the text sample set until the model training condition is satisfied, so as to obtain the text encoding model.

[0024] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0025] The processing module is specifically configured to, for each text sample, respectively encode the text sample twice through the text encoding model to be trained to obtain the original encoding vector and the positive example encoding vector corresponding to the text sample.

[0026] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0027] The processing module is specifically configured to, for each text sample, perform a synonym replacement process on at least one text unit in the text sample to obtain a target text sample corresponding to the text sample;

[0028] For each text sample, encode the text sample through the text encoding model to be trained to obtain the original encoding vector corresponding to the text sample;

[0029] For each text sample, encode the target text sample corresponding to the text sample through the text encoding model to be trained to obtain the positive example encoding vector corresponding to the text sample.

[0030] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0031] The processing module is specifically configured to, for each text sample, perform an addition and deletion process on the text sample to obtain a target text sample corresponding to the text sample, where the addition and deletion process includes deleting at least one text unit in the text sample, or adding at least one text unit to the text sample;

[0032] For each text sample, encode the text sample through the text encoding model to be trained to obtain the original encoding vector corresponding to the text sample;

[0033] For each text sample, encode the target text sample corresponding to the text sample through the text encoding model to be trained to obtain the positive example encoding vector corresponding to the text sample.

[0034] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0035] The processing module is specifically configured to, for each text sample, determine the cosine similarity between the original encoding vector corresponding to the text sample and the positive example encoding vector;

[0036] For each text sample, determine N cosine similarities between the original encoding vector corresponding to the text sample and N negative example encoding vectors;

[0037] For each text sample, calculate the loss value corresponding to the text sample according to the cosine similarity between the original encoding vector corresponding to the text sample and the positive example encoding vector, and the N cosine similarities between the original encoding vector corresponding to the text sample and the N negative example encoding vectors.

[0038] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0039] The acquisition module is further configured to acquire a text sample set, where the text sample set includes M sample triples, each sample triple includes a premise text sample, an inheritance text sample, and an opposing text sample, and M is an integer greater than or equal to 1;

[0040] The generation module is further configured to, for each sample triple, generate the original encoding vector corresponding to the premise text sample, the positive example encoding vector corresponding to the inheritance text sample, and the negative example encoding vector corresponding to the opposing text sample;

[0041] The processing module is further configured to, for each sample triple, calculate the loss value corresponding to the sample triple according to the original encoding vector corresponding to the premise text sample, the positive example encoding vector corresponding to the inheritance text sample, and the negative example encoding vector corresponding to the opposing text sample;

[0042] The training module is further configured to update the model parameters of the text encoding model to be trained according to the loss value corresponding to each sample triple in the text sample set until the model training condition is satisfied, and obtain the text encoding model.

[0043] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0044] The generation module is specifically configured to, for each sample triple, encode the premise text sample in the sample triple through the text encoding model to be trained to obtain the original encoding vector corresponding to the premise text sample;

[0045] For each sample triple, the inheritance text sample in the sample triple is encoded by the text encoding model to be trained, and the positive example encoding vector corresponding to the inheritance text sample is obtained;

[0046] For each sample triple, the opposing text sample in the sample triple is encoded by the text encoding model to be trained, and the negative example encoding vector corresponding to the opposing text sample is obtained.

[0047] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0048] The generation module is specifically configured to, for each sample triple, encode the premise text sample in the sample triple at least twice through the text encoding model to be trained, and obtain at least two original encoding vectors corresponding to the premise text sample;

[0049] For each sample triple, the inheritance text sample in the sample triple is encoded at least twice through the text encoding model to be trained, and at least two positive example encoding vectors corresponding to the inheritance text sample are obtained;

[0050] For each sample triple, the opposing text sample in the sample triple is encoded at least twice through the text encoding model to be trained, and at least two negative example encoding vectors corresponding to the opposing text sample are obtained.

[0051] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0052] The processing module is specifically configured to, for each sample triple, determine the first cosine similarity between the original encoding vector corresponding to the premise text sample and the positive example encoding vector corresponding to the inheritance text sample;

[0053] For each sample triple, determine M second cosine similarities between the original encoding vector corresponding to the premise text sample and the positive example encoding vectors corresponding to M inheritance text samples;

[0054] For each sample triple, determine M third cosine similarities between the original encoding vector corresponding to the premise text sample and the negative example encoding vectors corresponding to M opposing text samples;

[0055] For each sample triple, calculate the loss value corresponding to the sample triple according to the first cosine similarity, the M second cosine similarities, and the M third cosine similarities.

[0056] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0057] An acquisition module is further configured to acquire a text sample set, where the text sample set includes N text samples and M sample triples, each sample triple includes a premise text sample, a successor text sample, and an opposing text sample, N is an integer greater than 1, and M is an integer greater than or equal to 1;

[0058] A processing module is further configured to perform data augmentation processing on each text sample in the text sample set to obtain a positive example encoding vector corresponding to each text sample;

[0059] The acquisition module is further configured to acquire N negative example encoding vectors corresponding to each text sample in the text sample set, where the N negative example encoding vectors are generated based on the N text samples;

[0060] A training module is further configured to update the model parameters of the text encoding model to be trained according to the positive example encoding vector corresponding to each text sample and the N negative example encoding vectors to obtain a target text encoding model;

[0061] A generation module is further configured to generate, for each sample triple, an original encoding vector corresponding to the premise text sample, a positive example encoding vector corresponding to the successor text sample, and a negative example encoding vector corresponding to the opposing text sample;

[0062] The training module is further configured to fine-tune the target text encoding model according to the original encoding vector corresponding to the premise text sample, the positive example encoding vector corresponding to the successor text sample, and the negative example encoding vector corresponding to the opposing text sample in each sample triple to obtain a text encoding model.

[0063] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0064] The generation module is specifically configured to calculate the similarity between each abnormal text vector in the T abnormal text vectors and each text vector to be measured to obtain K similarities, where the greater the similarity, the higher the similarity degree between the abnormal text vector and the text vector to be measured;

[0065] For each abnormal text vector in the T abnormal text vectors, determine P text vectors to be measured with the greatest similarity to the abnormal text vector according to the K similarities, where P is an integer greater than or equal to 1 and less than K;

[0066] Determine the number of text vectors to be measured containing abnormal content among the K text vectors to be measured according to the P text vectors to be measured corresponding to each abnormal text vector;

[0067] Generate a detection result for the account to be detected according to the number of text vectors to be measured containing abnormal content.

[0068] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0069] The generation module is specifically configured to calculate the similarity between each of the K text vectors to be measured and each of the T abnormal text vectors respectively, to obtain K×T similarities. Among them, the greater the similarity, the higher the similarity degree between the abnormal text vector and the text vector to be measured;

[0070] If there is at least one similarity greater than or equal to the similarity threshold among the K×T similarities, then determine the number of text vectors to be measured containing abnormal content among the K text vectors to be measured according to the at least one similarity;

[0071] Generate a detection result for the account to be detected according to the number of text vectors to be measured containing abnormal content.

[0072] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0073] The generation module is specifically configured to initialize Q clustering centers, where Q is an integer greater than 1;

[0074] Calculate the Euclidean distance between each of the K text vectors to be measured and each of the T abnormal text vectors and each clustering center respectively, to obtain (K + T)×Q Euclidean distances;

[0075] According to the (K + T)×Q Euclidean distances, divide each text vector to be measured and each abnormal text vector into the clustering cluster corresponding to the clustering center with the smallest Euclidean distance respectively, to obtain Q clustering clusters;

[0076] According to the Q clustering clusters, divide each text vector to be measured and each abnormal text vector until the clustering optimization condition is satisfied, to obtain Q target clustering clusters;

[0077] Generate a detection result for the account to be detected according to the number of abnormal text vectors included in each target clustering cluster among the Q target clustering clusters.

[0078] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0079] The generation module is specifically configured to use the K text vectors to be measured and the T abnormal text vectors as a text vector set, where the text vector set includes (K + T) text vectors;

[0080] Select a text vector from the text vector set as the first text vector;

[0081] Calculate the similarity between the first text vector and the remaining (K + T - 1) text vectors in the text vector set respectively to obtain the second text vector corresponding to the maximum similarity;

[0082] If the maximum similarity is greater than or equal to the target threshold, add the second text vector to the first cluster corresponding to the first text vector and update the cluster center of the first cluster;

[0083] If the maximum similarity is less than the target threshold, use the second text vector as the cluster center of the second cluster;

[0084] When each text vector in the text vector set is divided into the corresponding cluster, R clusters are obtained, where R is an integer greater than or equal to 1;

[0085] Generate a detection result for the account to be detected according to the number of abnormal text vectors included in each of the R clusters.

[0086] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0087] The processing module is further configured to, after generating a detection result for the account to be detected according to the K text vectors to be measured and the T abnormal text vectors, intercept the text sent by the suspicious account when the detection result indicates that the account to be detected is a suspicious account;

[0088] The processing module is further configured to, when the detection result indicates that the account to be detected is a malicious account, intercept the text sent by the malicious account and ban the usage permission of the malicious account.

[0089] Another aspect of the present application provides a computer device, including a memory and a processor, where the memory stores a computer program, and the processor implements the methods of the above aspects when executing the computer program.

[0090] Another aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the computer program implements the methods of the above aspects when being executed by a processor.

[0091] Another aspect of the present application provides a computer program product, including a computer program, and the computer program implements the methods of the above aspects when being executed by a processor.

[0092] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0093] In an embodiment of the present application, a method for business processing is provided. First, a set of texts to be tested from a to-be-detected account is obtained. Then, each text to be tested in the set of texts to be tested is encoded using a text encoding model trained based on contrastive learning to obtain K text vectors to be tested. In addition, a set of abnormal texts also needs to be obtained, and then each abnormal text in the set of abnormal texts is encoded using this text encoding model to obtain T abnormal text vectors. Finally, a detection result for the to-be-detected account can be generated based on the K text vectors to be tested and the T abnormal text vectors. Through the above method, since contrastive learning does not need to focus on the cumbersome details of instances, but learns to distinguish data in the feature space at the abstract semantic level, the model optimization is simpler and the generalization ability is stronger. Therefore, using the text encoding model generated based on contrastive learning to encode the texts to be tested and the abnormal texts can obtain encoding vectors with better quality. Based on this, using higher-quality encoding vectors for text detection can improve the accuracy of text detection, and further improve the accuracy and reliability of account detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] Figure 1 It is an application schematic diagram of a greeting scenario in an embodiment of the present application;

[0095] Figure 2 It is an application schematic diagram of a voice friend adding scenario in an embodiment of the present application;

[0096] Figure 3 It is an application schematic diagram of a receiving new email scenario in an embodiment of the present application;

[0097] Figure 4 It is an implementation environment schematic diagram of a business processing method in an embodiment of the present application;

[0098] Figure 5 It is a process schematic diagram of a business processing method in an embodiment of the present application;

[0099] Figure 6 It is a schematic diagram of implementing account detection based on a text encoding model in an embodiment of the present application;

[0100] Figure 7 It is a schematic diagram of unsupervised contrastive learning in an embodiment of the present application;

[0101] Figure 8 It is a schematic diagram of supervised contrastive learning in an embodiment of the present application;

[0102] Figure 9 It is a schematic diagram of semi-supervised contrastive learning in an embodiment of the present application;

[0103] Figure 10A schematic diagram for processing a suspicious account in an embodiment of the present application;

[0104] Figure 11 A schematic diagram for processing a malicious account in an embodiment of the present application;

[0105] Figure 12 A schematic diagram of a service processing device in an embodiment of the present application;

[0106] Figure 13 A schematic structural diagram of a computer device in an embodiment of the present application. Detailed implementation manners

[0107] The embodiments of the present application provide a method, related device, equipment, and storage medium for service processing. By using a text encoding model generated based on contrastive learning, better-quality encoding vectors can be obtained, which is beneficial to improving the accuracy of text detection, and further improving the accuracy and reliability of account detection.

[0108] In the embodiments of the present application, the terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims, and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "correspond to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0109] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.

[0110] In the Internet age, with the rise of social media and online platforms, more and more people have begun to use various online names and avatars to show their identities. However, once these virtual identities are abused, they will bring a series of problems to the Internet platform, affecting the legitimate rights and interests of normal users and damaging the platform's ecological environment. Once the Internet platform identifies an abnormal account (for example, a suspicious account or a malicious account), it will deal with the abnormal account to avoid losses to normal users. In order to improve the efficiency of account detection, the text sent by the account can be encoded based on the artificial intelligence (AI) model to generate the corresponding text vector. The text vector can be used to detect whether the text involves abnormal content, and then identify abnormal accounts.

[0111] Most of the current mainstream text vector generation models use the bidirectional encoder representation from transformers (BERT) model. However, when the text volume is large, the BERT model will consume a lot of resources. In addition, although the BERT model has achieved good results in many downstream tasks, the generated text vector is not good. The main reason for this phenomenon is that the goal of BERT model pre-training does not consider how to generate a text vector suitable for vector search or similarity calculation. The text vector output by the BERT model is anisotropic and unevenly distributed in space. Therefore, the distance between text vectors cannot well represent the correlation between texts. Even if two texts are very similar in semantics, the distance between their previous text vectors may be large. Conversely, even if two texts are semantically far apart, the distance between their previous text vectors may be small. This phenomenon may lead to poor performance when using the BERT model for downstream tasks (for example, text clustering, similarity calculation, etc.).

[0112] Although the BERT model can capture contextual information in the text, its training goal is not to directly optimize the quality or similarity of text representation. The pre-training goal of BERT is to learn contextual representations through masked language model (MLM) and next sentence prediction (NSP) tasks. These tasks enable the BERT model to learn the structure and semantic information of the text, but do not guarantee that the generated representation vector has good performance in similarity measurement.

[0113] Based on this, in the embodiments of the present application, a method for business processing is provided. By introducing the method of contrastive learning, the quality of text representation is optimized, enabling it to perform better in similarity measurement and downstream tasks, and alleviating to a certain extent the problem of uneven and non-smooth vector distribution, thereby improving the accuracy and reliability of account detection. For the business processing method of the present application, when applied, it includes at least one of the following scenarios.

[0114] Scenario 1: Greeting scenario;

[0115] With the popularization of the Internet and mobile devices, social software has become an indispensable part of people's lives. People can use social software to establish connections with friends, family, colleagues, etc. at any time and place, sharing life, feelings, and information. Some social software supports searching for nearby people, chatting and greeting, and starting video or voice chats.

[0116] Exemplarily, a user can add a new user by searching for nearby people or entering relevant information (such as mobile phone number, social account, etc.). For ease of understanding, please refer to Figure 1 , Figure 1 which is an application schematic diagram of the greeting scenario in the embodiments of the present application. As shown in the figure, user B sends a friend addition request to user A by searching for nearby friends and sends the greeting text indicated by A1 (for example, "What a coincidence. We seem to be junior high school classmates. Add me.") to user A. Based on this, the method provided by the present application can be used to encode the greeting text to obtain the corresponding text vector. Based on the text vector, it can be further determined whether the account used by user B is abnormal.

[0117] Scenario 2: Voice friend addition scenario;

[0118] Social software based on voice matching supports users to match friends online by voice. Taking random matching as an example, a user can become friends with users from different provinces and different cities.

[0119] Exemplarily, a user can add a new user by searching for nearby people or random matching. For ease of understanding, please refer to Figure 2 , Figure 2This is an application schematic diagram of the voice friend adding scenario in the embodiments of the present application. As shown in the figure, user B initiates a friend adding application to user A through random matching and sends the voice indicated by B1 to user A. Then, based on automatic speech recognition (ASR) technology, the voice is converted into text (for example, "Your voice sounds like my junior high school classmate. Add me."). Based on this, the method provided by the present application can be used to encode the text obtained after voice conversion to obtain a corresponding text vector. Based on the text vector, it is further determined whether there is an abnormality in the account used by user B.

[0120] Scenario 3: Receiving a new email scenario;

[0121] Using the email service is an information transfer service based on computers and communication networks, which can provide users with various types of information such as transmitting electronic letters, documents, images, and digital voices. Email enables people to receive and send letters anywhere and at any time, solving the limitations of time and space and greatly improving work efficiency.

[0122] Exemplarily, a complete email consists of two basic parts: the header and the body. Among them, the header includes the recipient's email address and the subject. The subject text is a general description of the email content, which can be a word or a sentence. It is drafted by the sender. For ease of understanding, please refer to Figure 3 , Figure 3 This is an application schematic diagram of the receiving new email scenario in the embodiments of the present application. As shown in the figure, C1 is used to indicate the subject text of a certain email received by the user (that is, "Congratulations on winning the reward sent by our company. Please claim it as soon as possible"). Based on this, the method provided by the present application can be used to encode the subject text to obtain a corresponding text vector. Based on the text vector, it is further determined whether there is an abnormality in the account used by the user who sent the email.

[0123] It should be noted that the above application scenarios are only examples, and the service processing method provided in this embodiment can also be applied to other scenarios, which are not limited here.

[0124] It can be understood that the solution provided in the embodiments of the present application involves technologies such as natural language processing (NLP) and machine learning (ML) of AI.

[0125] Using contrastive learning in NLP scenarios can optimize the quality of text representation. Among them, NLP is an important direction in the fields of computer science and AI. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. NLP involves natural language, that is, the language people use in daily life, and is closely related to linguistics research; at the same time, it involves computer science and mathematics. The pre-training model, an important technology for model training in the AI field, has evolved from the large language model in the NLP field. After fine-tuning, the large language model can be widely applied to downstream tasks. NLP technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs and other technologies.

[0126] Among them, the pre-training model, also known as the foundation model or large model, refers to a deep neural network (DNN) with a large number of parameters. It is trained on a large amount of unlabeled data, and uses the function approximation ability of the large-parameter DNN to enable the pre-training model to extract common features from the data. Through techniques such as fine-tuning, parameter-efficient fine-tuning (PEFT), and prompt-tuning, it is suitable for downstream tasks. Therefore, the pre-training model can achieve ideal results in few-shot or zero-shot scenarios. The pre-training model can be divided into language models, visual models, speech models, multi-modal models, etc. according to the data modalities processed.

[0127] The text encoding model involved in this application can be learned and trained based on ML technology. Among them, ML is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. ML is the core of AI and the fundamental way to make computers intelligent. Its applications cover all fields of AI. ML and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The pre-training model is the latest development result of deep learning, integrating the above technologies.

[0128] The method provided in this application can be applied to Figure 4The illustrated implementation environment includes a first terminal 110, a second terminal 120, and a server 130. Moreover, each terminal and the server 130 can communicate through a communication network 140. Among them, the communication network 140 uses standard communication technologies and / or protocols, usually the Internet, but can also be any network, including but not limited to any combination of Bluetooth, local area network (LAN), metropolitan area network (MAN), wide area network (WAN), mobile, private network, or virtual private network. In some embodiments, customized or dedicated data communication technologies can be used to replace or supplement the above data communication technologies.

[0129] The terminals involved in this application include but are not limited to mobile phones, tablets, laptops, desktop computers, intelligent voice interaction devices, virtual reality devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. Among them, the client is deployed on the terminal. The client can run on the terminal in the form of a browser, or can also run on the terminal in the form of an independent application (APP) or a small program, etc.

[0130] The server 130 involved in this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and AI platforms. It should be noted that this application takes the configuration and deployment of the text encoding model on the server 130 as an example for illustration. In some embodiments, the configuration of the text encoding model can also be deployed on the terminal (for example, the second terminal 120). In some embodiments, part of the configuration of the text encoding model is deployed on the terminal, and part of the configuration is deployed on the server 130.

[0131] Combined with the above implementation environment, in step S1, the sender uses the first terminal 110 to send a text (e.g., a greeting text), which serves as the text to be tested, and the first terminal 110 sends it to the server 130 through the communication network 140. In step S2, the server 130 calls the text encoding model to encode the K texts to be tested, obtaining K text vectors to be tested, where the K texts to be tested at least include the text to be tested sent by the first terminal 110. In step S3, the server 130 calls the text encoding model to encode the T abnormal texts in the black library, obtaining T abnormal text vectors. In step S4, the server 130 generates a detection result for the account to be detected based on the K text vectors to be tested and the T abnormal text vectors, where the account to be detected is the account used by the sender.

[0132] In step S5, the server 130 feeds back the account result of the account to be detected to the first terminal 110 through the communication network 140. If the detection result indicates that the account to be detected is a suspicious account, the text to be tested is intercepted. If the detection result indicates that the account to be detected is a malicious account, the malicious account is blocked. If the detection result indicates that the account to be detected is a normal account, then in step S6, the server 130 sends the text to be tested to the second terminal 120 used by the recipient through the communication network 140.

[0133] In view of the fact that this application involves some terms related to professional fields, for the sake of easy understanding, explanations will be given below.

[0134] (1) Simple contrastive learning of sentence embeddings (SimCSE): Based on the idea of contrastive learning, the distance between text vectors representing the same semantics is close enough, and the distance between text vectors representing different semantics is as far apart as possible. Whether SimCSE uses unsupervised learning or supervised learning, it has achieved excellent model effects.

[0135] (2) Malicious registration: A registration behavior aimed at implementing acts violating national laws and regulations, or a registration behavior aimed at violating the account usage agreement.

[0136] (3) Account blocking: Punishing a malicious account, forcing the account to go offline, and prohibiting login.

[0137] (4) Normal account: An account with a relatively long active time, having regular social activities (e.g., interacting with friends via messages, posting and commenting on the Moments, reading official account articles, etc.), and having no punishment record.

[0138] (5) Malicious Account: An account that is used for the purpose of committing acts in violation of national laws and regulations, or an account that is used for the purpose of violating the account usage agreement.

[0139] (6) Suspicious Account: An account selected through certain rules, including both real malicious accounts and normal accounts whose behaviors are similar to those of malicious accounts but do not belong to malicious accounts.

[0140] (7) Greeting Text: Refers to the text content of the friend addition application sent when the APP adds a friend.

[0141] (8) Blacklist: The evidence content of user complaints, or the malicious greeting text content determined after manual customer service review.

[0142] Combined with the above introduction, the business processing method in this application will be introduced below. Please refer to Figure 5 , the business processing method in the embodiments of this application can be completed independently by the server, can also be completed independently by the terminal, or can be completed in cooperation between the terminal and the server. The method of this application includes:

[0143] 210. Obtain a set of texts to be tested, where the set of texts to be tested includes K texts to be tested sent by the account to be detected, and K is an integer greater than or equal to 1;

[0144] In one or more embodiments, in one case, obtain the texts (such as greeting texts) sent by all accounts within a period of time (for example, 2 hours), and then select the texts sent by the account to be detected from these texts sent by the accounts as K texts to be tested. At this time, K≥1. In one case, the texts sent by the account to be detected within the historical time period can be directly used as K texts to be tested. At this time, K≥1. In one case, the text sent by the account to be detected most recently can be used as the text to be tested. At this time, K = 1.

[0145] 220. Encode each text to be tested in the set of texts to be tested using a text encoding model to obtain K text vectors to be tested, where the K text vectors to be tested have a one-to-one correspondence with the K texts to be tested, and the text encoding model is obtained by contrastive learning using positive example encoding vector pairs and negative example encoding vector pairs;

[0146] In one or more embodiments, call the text encoding model to encode each text to be tested in the set of texts to be tested, so as to obtain the text vectors to be tested corresponding to each text to be tested, that is, obtain K text vectors to be tested.

[0147] Specifically, the text encoding model is a model generated after contrastive learning (e.g., SimCSE) using positive example encoding vector pairs and negative example encoding vector pairs. It should be noted that the text encoding model can specifically adopt a BERT model structure, a residual network (ResNet) structure, a robustly optimized BERT (RoBERTa) model, etc., which is not limited here.

[0148] 230. Obtain an abnormal text set, where the abnormal text set includes T abnormal texts that have been labeled as abnormal content, and T is an integer greater than or equal to 1.

[0149] In one or more embodiments, obtain the abnormal text set from a black library, where each abnormal text included in the abnormal text set has been labeled as abnormal content.

[0150] Specifically, users can file complaints about the received text. Then, these complained texts need to go through data cleaning. One way is manual cleaning, that is, customer service personnel check the complained texts and add the texts involving abnormal content to the black library. Another way is automatic cleaning, that is, detect whether the complained texts include words in the blacklist, and add the complained texts including words in the blacklist to the black library.

[0151] 240. Use the text encoding model to encode each abnormal text in the abnormal text set to obtain T abnormal text vectors, where the T abnormal text vectors have a one-to-one correspondence with the T abnormal texts.

[0152] In one or more embodiments, call the text encoding model to encode each abnormal text in the abnormal text set, thereby obtaining the abnormal text vector corresponding to each abnormal text, that is, obtaining T abnormal text vectors.

[0153] 250. Generate a detection result for the account to be detected according to K text vectors to be measured and T abnormal text vectors, where the detection result is used to represent the abnormal degree of the account to be detected.

[0154] In one or more embodiments, based on the K text vectors to be measured and the T abnormal text vectors, similarity calculation or clustering can be performed to determine the detection result (i.e., whether it belongs to abnormal text) of each text to be measured. According to the detection result of each text to be measured, further generate the detection result of the account to be detected (i.e., the abnormal degree of the account to be detected).

[0155] Specifically, for ease of understanding, please refer to Figure 6 , Figure 6This is a schematic diagram of implementing account detection based on a text encoding model in an embodiment of the present application. As shown in the figure, first, data needs to be collected and a data set needs to be constructed, wherein the data set includes texts (e.g., greeting texts) sent by users every day (or every hour, etc.) and confirmed abnormal texts (i.e., texts obtained after data cleaning based on the complained texts). Then, the set of texts to be tested sent by the account to be tested and the set of abnormal texts in the black library are extracted. The text encoding model is then called to encode the K texts to be tested in the set of texts to be tested, respectively, to obtain K text vectors to be tested, and the text encoding model is called to encode the T abnormal texts in the set of abnormal texts, respectively, to obtain T abnormal text vectors. Based on this, similarity calculation or clustering is performed based on the K text vectors to be tested and the T abnormal text vectors to determine the detection result of the account to be detected. Finally, corresponding disposal is performed based on the detection result of the account to be detected.

[0156] Since the underlying text detection service strongly relies on the algorithm for generating text vectors, the higher the quality of the generated text vectors, the better the similarity recall results.

[0157] In an embodiment of the present application, a method for business processing is provided. In the above manner, the text to be tested and the abnormal text are encoded using a text encoding model generated based on contrastive learning, which can semantically measure whether the two texts are similar, has better globality, and thus can obtain a better quality encoding vector. Based on this, using a higher quality encoding vector for text detection can improve the accuracy of text detection, thereby improving the accuracy and reliability of account detection (it was found on the business landing side that the business gain can be increased by about 8%). In addition, the introduction of a model for text detection has stronger robustness, which can better prevent malicious users from bypassing real-time policies through batch testing.

[0158] Optionally, in the above Figure 5 Based on the corresponding one or more embodiments, another optional embodiment provided by the embodiment of the present application may also include:

[0159] Get a text sample set, where the text sample set includes N text samples, where N is an integer greater than 1;

[0160] Perform data enhancement processing on each text sample in the text sample set to obtain the positive example encoding vector corresponding to each text sample;

[0161] Obtain N negative example encoding vectors corresponding to each text sample in the text sample set, where the N negative example encoding vectors are generated based on the N text samples;

[0162] For each text sample, a loss value corresponding to the text sample is calculated according to the positive example encoding vector corresponding to the text sample and N negative example encoding vectors.

[0163] According to the loss value corresponding to each text sample in the text sample set, the model parameters of the text encoding model to be trained are updated until the model training conditions are met, and a text encoding model is obtained.

[0164] In one or more embodiments, a method for generating a text encoding model based on unsupervised contrast learning is introduced. As can be seen from the foregoing embodiments, a text sample set for training is obtained, and the text sample set includes at least two text samples. Among them, the text sample set can be understood as a batch of data. In unsupervised contrast learning, the model does not require paired similar text samples as training data, but generates positive example encoding vectors (i.e., identifiers of similar text samples) by performing data augmentation processing on the same text sample.

[0165] Specifically, for ease of understanding, please refer to Figure 7 , Figure 7 is a schematic diagram of unsupervised contrast learning in the embodiments of the present application. As shown in the figure, taking a certain text sample in the text sample set as an example, assume that the text sample is "Two little dogs are running". Based on this, data augmentation processing is performed on the text sample to generate a positive example encoding vector corresponding to the text sample. At the same time, the encoding vectors obtained by encoding other text samples are used as negative example encoding vectors for this text sample (i.e., "Two little dogs are running").

[0166] Exemplarily, the text sample "Two little dogs are running" and the text sample "A person is surfing on the sea" form a pair of negative samples. Therefore, the encoding vector obtained by encoding the text sample "A person is surfing on the sea" is a negative example encoding vector for the text sample "Two little dogs are running".

[0167] Based on this, the loss value corresponding to each text sample can be calculated, and the N loss values are summed to obtain a target loss value. Thus, the model parameters of the text encoding model to be trained are optimized by the methods of backpropagation and gradient descent. Until the model training conditions are met, a text encoding model is obtained.

[0168] It should be noted that in one case, an exhaustion criterion can be used as the basis for determining whether the model training conditions are met. For example, an iteration number threshold is set, and when the iteration number reaches the iteration number threshold, it means that the model training conditions have been met. In another case, an observation criterion can be used as the basis for determining whether the model training conditions are met. For example, when the loss value has converged, it means that the model training conditions have been met.

[0169] Secondly, in the embodiments of the present application, a method for generating a text encoding model based on unsupervised contrastive learning is provided. Through the above method, considering that there may be a lack of a large amount of labeled sample data in practical applications (that is, to label whether two sample texts are similar texts), therefore, training the model based on unsupervised learning can solve the problem of a small number of labeled samples. It can not only save the cost of manual annotation, but also learn features with better adaptability and richness.

[0170] Optionally, based on the above Figure 5 corresponding one or more embodiments, in another optional embodiment provided by the embodiments of the present application, data augmentation processing is performed on each text sample in the text sample set to obtain the positive example encoding vector corresponding to each text sample, which may specifically include:

[0171] For each text sample, the text sample is encoded twice by the text encoding model to be trained to obtain the original encoding vector and the positive example encoding vector corresponding to the text sample.

[0172] In one or more embodiments, a method for data augmentation based on a dropout mask is introduced. As can be seen from the foregoing embodiments, considering that if two exactly the same text samples are used as a positive example pair, the generalization ability will be poor. Therefore, some data augmentation means are needed to make the two text samples in the positive example pair different. Hereinafter, dropout mask will be used as a data augmentation method to generate positive example pairs.

[0173] Specifically, assume that the text sample set is represented as Then a copy of the text sample set is obtained, and the dropout mask is used for the two text sample sets respectively. The two dropout mask operations are defined as z and z' respectively. Taking the i-th text sample as an example, this text sample is represented as x in the text sample set i . Based on this, the text sample x i is used as the input of the text encoding model to be trained (f θ ), and the original encoding vector corresponding to the text sample x i is output through the text encoding model to be trained That is, Again, the text sample x i is used as the input of the text encoding model to be trained, and the positive example encoding vector corresponding to the text sample x i is output through the text encoding model to be trained That is,

[0174] SimCSE uses the same text samples as input, but for each input, a different dropout mask is used. Dropout is a regularization technique commonly used to prevent neural networks from overfitting. During training, dropout randomly discards the outputs of some neurons, preventing the model from relying too much on individual neurons. In SimCSE, due to the randomness of dropout, slightly different encoded vectors are generated each time the same text sample is input. These representations are considered positive example pairs because they all come from the same text sample and are thus semantically similar.

[0175] Dropout provides a smoother and more stable data augmentation method. Random dropout is a continuous data augmentation method as it proportionally reduces neuron outputs rather than completely deleting them. This smoothness enables the model to better learn the subtle differences in sentence representations. Random dropout forces the model not to rely too much on individual neurons during training, thereby improving the model's robustness. By learning similar sentence representations under different dropout scenarios, the model can better generalize to new or unseen data. Random dropout can maintain a constant dropout rate throughout the training process, making model training more stable. Random dropout can be applied to various neural network architectures and tasks.

[0176] Again, in the embodiments of the present application, a way to implement data augmentation based on dropout mask is provided. Through the above method, due to the randomness of random inactivation, slightly different encoded vectors are output when the same text sample is input. Since they come from the same text sample, they are semantically similar and can be directly used as a set of positive example pairs. Thus, the problem of insufficient data annotation samples can be effectively solved without manual data annotation, thereby improving the efficiency and convenience of model training.

[0177] Optionally, based on one or more corresponding embodiments above, in another optional embodiment provided by the embodiments of the present application, data augmentation processing is performed on each text sample in the text sample set to obtain the positive example encoded vector corresponding to each text sample, which may specifically include: Figure 5 For each text sample, perform synonym replacement processing on at least one text unit in the text sample to obtain the target text sample corresponding to the text sample;

[0178] For each text sample, encode the text sample through the text encoding model to be trained to obtain the original encoded vector corresponding to the text sample;

[0179]

[0180] ​For each text sample, the target text sample corresponding to the text sample is encoded by the text encoding model to be trained, and the positive example encoding vector corresponding to the text sample is obtained.

[0181] In one or more embodiments, a method for data augmentation based on synonym replacement is introduced. As can be seen from the foregoing embodiments, considering that if two completely identical text samples are used as a positive example pair, the generalization ability will be poor. Therefore, some data augmentation means are needed to make the two text samples of the positive example pair different. Below, synonym replacement processing will be used as a data augmentation method to generate positive example pairs.

[0182] Specifically, assume that the text sample set is represented as Then, synonym replacement processing is performed on each text sample in the text sample set, and another text sample set is obtained therefrom Exemplarily, a text sample consists of at least one text unit. Assume the text sample is "Two little dogs are running". Among them, the text unit "running" is replaced with the synonym "galloping", and the target text sample corresponding to the text sample is "Two little dogs are galloping".

[0183] It can be understood that if the text sample is in Chinese, the text unit is a Chinese character or word. If the text sample is in English, the text unit is a word. In addition, the text sample can also be in other languages. This is only an illustration here and should not be construed as a limitation to this application.

[0184] Thus, the text sample set and the text sample set are encoded respectively. Taking the i-th text sample as an example, the i-th text sample is represented as x in the text sample set i , and the target text sample corresponding to the i-th text sample is represented as in the text sample set Based on this, the text sample x i is used as the input of the text encoding model to be trained (f θ ), and the original encoding vector i corresponding to the text sample x is output through the text encoding model to be trained, that is, The target text sample i corresponding to the text sample x is used as the input of the text encoding model to be trained, and the positive example encoding vector i corresponding to the text sample x is output through the text encoding model to be trained, that is,

[0185] Again, in the embodiments of the present application, a method for data augmentation based on synonym replacement is provided. Through the above method, the text units in the text sample can be replaced with synonyms, thereby achieving the effect of data amplification without changing the semantics. Thus, the performance and generalization ability of the model can be further improved.

[0186] Optionally, based on one or more corresponding embodiments above, in another optional embodiment provided by the embodiments of the present application, data augmentation processing is performed on each text sample in the text sample set to obtain the positive example encoding vector corresponding to each text sample, which may specifically include: Figure 5 For each text sample, addition and deletion processing is performed on the text sample to obtain the target text sample corresponding to the text sample, where the addition and deletion processing includes deleting at least one text unit in the text sample, or adding at least one text unit to the text sample;

[0187] For each text sample, the text sample is encoded by the text encoding model to be trained to obtain the original encoding vector corresponding to the text sample;

[0188] For each text sample, the target text sample corresponding to the text sample is encoded by the text encoding model to be trained to obtain the positive example encoding vector corresponding to the text sample.

[0189] In one or more embodiments, a method for data augmentation based on adding noise is introduced. As can be seen from the foregoing embodiments, considering that if two completely identical text samples are used as positive example pairs, the generalization ability will be poor. Therefore, some data augmentation means are needed to make the two text samples of the positive example pair different. Hereinafter, adding noise (i.e., addition and deletion processing of text samples) will be used as a data augmentation method to generate positive example pairs.

[0190] Specifically, assume that the text sample set is represented as

[0191] Then, addition and deletion processing is performed on each text sample in the text sample set, thereby obtaining another text sample set Exemplarily, a text sample is composed of at least one text unit. Assume that the text sample is "Two little dogs are running". Among them, by deleting the text unit "are" in the text sample, the target text sample corresponding to the text sample is "Two little dogs running". Exemplarily, assume that the text sample is "Two little dogs are running". Among them, by adding the text unit "in" to the text sample, the target text sample corresponding to the text sample is "Two little dogs are running in".

[0192] ​It can be understood that if the text sample is in Chinese, the text unit is a Chinese character or punctuation mark. If the text sample is in English, the text unit is a word or punctuation mark. In addition, the text sample can also be in other languages. This is only an illustration here and should not be construed as a limitation of this application.

[0193] Then, the text sample sets and the text sample sets are encoded respectively. Taking the i-th text sample as an example, the i-th text sample is represented as x in the text sample set i , and the target text sample corresponding to the i-th text sample is represented as in the text sample set Based on this, the text sample x i is used as the input of the text encoding model to be trained (f θ ), and the original encoding vector corresponding to the text sample x i is output through the text encoding model to be trained That is, The target text sample corresponding to the text sample x i is used as the input of the text encoding model to be trained, and the positive example encoding vector corresponding to the text sample x is output through the text encoding model to be trained i That is, That is,

[0194] Furthermore, in the embodiments of this application, a method for data augmentation based on adding noise is provided. Through the above method, the text units in the text sample can be deleted, or text units can be added to the text sample. Thus, the effect of data amplification can be achieved without changing the semantics. Thereby, the performance and generalization ability of the model can be further improved.

[0195] Optionally, based on one or more corresponding embodiments above, in another optional embodiment provided by the embodiments of this application, for each text sample, a loss value corresponding to the text sample is calculated according to the positive example encoding vector and N negative example encoding vectors corresponding to the text sample. Specifically, it may include: Figure 5 For each text sample, determine the cosine similarity between the original encoding vector corresponding to the text sample and the positive example encoding vector;

[0196] For each text sample, determine the cosine similarity between the original encoding vector corresponding to the text sample and the N negative example encoding vectors;

[0197] For each text sample, determine N cosine similarities between the original encoding vector corresponding to the text sample and the N negative example encoding vectors;

[0198] For each text sample, a loss value corresponding to the text sample is calculated based on the cosine similarity between the original encoding vector corresponding to the text sample and the positive example encoding vector, and the N cosine similarities between the original encoding vector corresponding to the text sample and the N negative example encoding vectors.

[0199] In one or more embodiments, a method for calculating a loss value based on an unsupervised loss function is introduced. As can be seen from the foregoing embodiments, taking the text encoding model generated based on unsupervised contrastive learning as an example, its core idea is that after obtaining the original encoding vector and the positive example encoding vector corresponding to each text sample, N negative example encoding vectors corresponding to each text sample can be constructed respectively.

[0200] Specifically, the following method can be used to calculate the loss value corresponding to the i-th text sample:

[0201]

[0202] where L Unsup-SimCSE_i represents the loss value corresponding to the i-th text sample. represents the original encoding vector corresponding to the i-th text sample. represents the positive example encoding vector corresponding to the i-th text sample. represents the cosine similarity between the original encoding vector corresponding to the i-th text sample and the positive example encoding vector. represents the positive example encoding vector corresponding to the j-th text sample, that is, the negative example encoding vector corresponding to the i-th text sample. represents the cosine similarity between the original encoding vector corresponding to the i-th text sample and a negative example encoding vector. Since the value range of j is from 1 to N, N cosine similarities can be obtained. N represents the number of text samples included in the text sample set. τ represents the temperature coefficient. For example, it can be taken as 0.05.

[0203] sim(·) in formula (1) represents similarity calculation. Taking the cosine similarity as an example, that is, the following method is used to calculate the similarity between two encoding vectors:

[0204]

[0205] where h 1 represents an encoding vector, and h 2 represents another encoding vector.

[0206] Based on formula (1), the following method can be used to sum the loss values corresponding to N text samples, that is:

[0207]

[0208] Among them, L Unsup-SimCSE represents the target loss value.

[0209] Based on this, the model parameters of the text encoding model to be trained are optimized using the target loss value until the model training conditions are met, and a text encoding model is obtained.

[0210] Furthermore, in the embodiments of the present application, a method for calculating the loss value based on an unsupervised loss function is provided. Through the above method, the contrastive learning loss function is used to minimize the distance between positive example pairs and maximize the distance from negative example pairs. Thereby, the model is enabled to learn to generate text vectors with good similarity metrics.

[0211] Optionally, based on one or more of the above Figure 5 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, it may further include:

[0212] Obtain a text sample set, where the text sample set includes M sample triples, and each sample triple includes a premise text sample, an inheritance text sample, and an opposing text sample, and M is an integer greater than or equal to 1;

[0213] For each sample triple, generate an original encoding vector corresponding to the premise text sample, a positive example encoding vector corresponding to the inheritance text sample, and a negative example encoding vector corresponding to the opposing text sample;

[0214] For each sample triple, calculate the loss value corresponding to the sample triple according to the original encoding vector corresponding to the premise text sample, the positive example encoding vector corresponding to the inheritance text sample, and the negative example encoding vector corresponding to the opposing text sample;

[0215] Update the model parameters of the text encoding model to be trained according to the loss value corresponding to each sample triple in the text sample set until the model training conditions are met, and a text encoding model is obtained.

[0216] In one or more embodiments, a method for generating a text encoding model based on supervised contrastive learning is introduced. As can be seen from the foregoing embodiments, a text sample set for training is obtained, and the text sample set includes at least one sample triple. Among them, the text sample set can be understood as a batch of data. In supervised contrastive learning, the model requires paired similar text samples as positive example pairs and forms negative example pairs with other text samples.

[0217] Specifically, for ease of understanding, please refer to Figure 8 , Figure 8This is a schematic diagram of supervised contrastive learning in an embodiment of the present application. As shown in the figure, taking a sample triple in a text sample set as an example, assume that the premise text sample is "Two little dogs are running", its corresponding entailment text sample is "There are two animals outdoors", and its corresponding contradiction text sample is "The pet is sitting on the bench". Based on this, the original encoding vector corresponding to the premise text sample, the positive example encoding vector corresponding to the entailment text sample, and the negative example encoding vector corresponding to the contradiction text sample can be obtained through the text encoding model to be trained. Then, according to the original encoding vector corresponding to the premise text sample, the positive example encoding vector corresponding to the entailment text sample, and the negative example encoding vector corresponding to the contradiction text sample in each sample triple, the loss value corresponding to each sample triple is calculated, and the M loss values are summed to obtain the target loss value. Thus, the model parameters of the text encoding model to be trained are optimized by the method of backpropagation and gradient descent. Until the model training conditions are met, a text encoding model is obtained.

[0218] It should be noted that there is the same or similar semantics between the premise text sample and the entailment text sample, and there is different semantics between the premise text sample and the contradiction text sample.

[0219] Secondly, in an embodiment of the present application, a method for generating a text encoding model based on supervised contrastive learning is provided. Through the above method, the model can also use paired premise text samples and entailment text samples for data augmentation. This is because these texts are semantically similar but may be different in expression. Therefore, paired similar sentences can provide richer information to help the model learn better text representations.

[0220] Optionally, on the basis of one or more corresponding embodiments described above, in another optional embodiment provided by the embodiment of the present application, for each sample triple, generating the original encoding vector corresponding to the premise text sample, the positive example encoding vector corresponding to the entailment text sample, and the negative example encoding vector corresponding to the contradiction text sample may specifically include: Figure 5 For each sample triple, encoding the premise text sample in the sample triple through the text encoding model to be trained to obtain the original encoding vector corresponding to the premise text sample;

[0221] For each sample triple, encoding the entailment text sample in the sample triple through the text encoding model to be trained to obtain the positive example encoding vector corresponding to the entailment text sample;

[0222] For each sample triple, encoding the contradiction text sample in the sample triple through the text encoding model to be trained to obtain the negative example encoding vector corresponding to the contradiction text sample;

[0223] For each sample triple, the opposing text sample in the sample triple is encoded by the text encoding model to be trained, and the negative example encoding vector corresponding to the opposing text sample is obtained.

[0224] In one or more embodiments, a method for generating an encoding vector is introduced. As can be seen from the foregoing embodiments, the text sample set includes M sample triples, that is, it is expressed as where x i represents the premise text sample in the i-th sample triple, represents the inheritance text sample in the i-th sample triple, represents the opposing text sample in the i-th sample triple.

[0225] Specifically, taking the i-th sample triple as an example, this sample triple is represented in the text sample set as Based on this, the premise text sample x i is used as the input of the text encoding model to be trained (f θ ), and the original encoding vector h i corresponding to the premise text sample x i is output through the text encoding model to be trained, that is, h i = f θ (x i ). The inheritance text sample is used as the input of the text encoding model to be trained, and the positive example encoding vector corresponding to the inheritance text sample is output through the text encoding model to be trained, that is, The opposing text sample is used as the input of the text encoding model to be trained, and the negative example encoding vector corresponding to the opposing text sample is output through the text encoding model to be trained, that is,

[0226] Again, in the embodiments of the present application, a method for generating an encoding vector is provided. Through the above method, since the sample triple includes the labeled premise text sample, inheritance text sample, and opposing text sample, the text encoding model to be trained can be directly called for encoding to obtain the corresponding encoding vector. Therefore, no additional operations are required, thereby improving the efficiency and effect of model training.

[0227] Optionally, in the above Figure 5Based on one or more corresponding embodiments, in another alternative embodiment provided by the embodiments of the present application, for each sample triple, an original encoding vector corresponding to the premise text sample, a positive example encoding vector corresponding to the inheritance text sample, and a negative example encoding vector corresponding to the opposing text sample are generated. Specifically, it may include:

[0228] For each sample triple, the premise text sample in the sample triple is encoded at least twice through the text encoding model to be trained, and at least two original encoding vectors corresponding to the premise text sample are obtained;

[0229] For each sample triple, the inheritance text sample in the sample triple is encoded at least twice through the text encoding model to be trained, and at least two positive example encoding vectors corresponding to the inheritance text sample are obtained;

[0230] For each sample triple, the opposing text sample in the sample triple is encoded at least twice through the text encoding model to be trained, and at least two negative example encoding vectors corresponding to the opposing text sample are obtained.

[0231] In one or more embodiments, another way of generating encoding vectors is introduced. As can be seen from the foregoing embodiments, the text sample set includes M sample triples, that is, it is expressed as where x i represents the premise text sample in the i-th sample triple, represents the inheritance text sample in the i-th sample triple, represents the opposing text sample in the i-th sample triple.

[0232] Specifically, taking the i-th sample triple as an example, this sample triple is represented in the text sample set as Then at least one copy of the sample triple is obtained, and each sample triple is used with dropout mask. Taking the use of two dropout masks as an example, the operations of the two dropout masks are defined as z and z' respectively. Based on this, the text sample x i is used as the input of the text encoding model to be trained (f θ ), and through the text encoding model to be trained, an original encoding vector corresponding to the text sample x i is output, that is, That is, The text sample x i is used as the input of the text encoding model to be trained again, and through the text encoding model to be trained, a positive example encoding vector corresponding to the text sample x i is output, that is, That is,

[0233] The inherited text sample is used as the input of the text encoding model to be trained, and the inherited text sample is output through the text encoding model to be trained corresponding to a positive example encoding vector That is The inherited text sample is used as the input of the text encoding model to be trained again, and the inherited text sample is output through the text encoding model to be trained corresponding to the positive example encoding vector That is

[0234] The opposing text sample is used as the input of the text encoding model to be trained, and the opposing text sample is output through the text encoding model to be trained corresponding to a negative example encoding vector That is The opposing text sample is used as the input of the text encoding model to be trained again, and the opposing text sample is output through the text encoding model to be trained corresponding to the negative example encoding vector That is

[0235] Thus, after obtaining the original encoding vector and the positive example encoding vector positive example encoding vector and the positive example encoding vector negative example encoding vector and the negative example encoding vector it is possible to form different triples for model training. For example For another example etc., thereby increasing the way of text representation and improving the generalization ability of the model

[0236] Again, in the embodiments of the present application, another way of generating encoding vectors is provided. Through the above method, for the precondition text sample, inherited text sample, and opposing text sample that have been annotated, dropout masks can be used to generate slightly different text representations, thereby helping the model learn more robust text representations and improving the generalization ability

[0237] Optionally, based on one or more corresponding embodiments described above, in another optional embodiment provided by the embodiments of the present application, for each sample triple, according to the original encoding vector corresponding to the precondition text sample, the positive example encoding vector corresponding to the inherited text sample, and the negative example encoding vector corresponding to the opposing text sample, the loss value corresponding to the sample triple is calculated, which may specifically include Figure 5 That is

[0238] For each sample triple, determine the first cosine similarity between the original encoding vector corresponding to the premise text sample and the positive example encoding vector corresponding to the inherited text sample;

[0239] For each sample triple, determine M second cosine similarities between the original encoding vector corresponding to the premise text sample and the positive example encoding vectors corresponding to M inherited text samples;

[0240] For each sample triple, determine M third cosine similarities between the original encoding vector corresponding to the premise text sample and the negative example encoding vectors corresponding to M adversarial text samples;

[0241] For each sample triple, calculate the loss value corresponding to the sample triple based on the first cosine similarity, M second cosine similarities, and M third cosine similarities.

[0242] In one or more embodiments, a method for calculating a loss value based on a supervised loss function is introduced. As can be seen from the foregoing embodiments, taking the text encoding model generated based on supervised contrastive learning as an example, the original encoding vector corresponding to the premise text sample and the positive example encoding vector corresponding to the inherited text sample are used as a positive example pair, and the original encoding vector corresponding to the premise text sample and the negative example encoding vector corresponding to other adversarial text samples are used as a negative example pair for contrastive learning.

[0243] Specifically, the following method can be used to calculate the loss value corresponding to the i-th text sample:

[0244]

[0245] where L Sup-SimCSE_i represents the loss value corresponding to the i-th text sample. h i represents the original encoding vector corresponding to the i-th premise text sample. represents the positive example encoding vector corresponding to the i-th inherited text sample. represents the first cosine similarity between the original encoding vector corresponding to the i-th premise text sample and the positive example encoding vector corresponding to the i-th inherited text sample. represents the positive example encoding vector corresponding to the j-th inherited text sample. represents the second cosine similarity between the original encoding vector corresponding to the i-th premise text sample and the positive example encoding vector corresponding to the j-th inherited text sample. Since the value range of j is from 1 to M, M second cosine similarities can be obtained. represents the negative example encoding vector corresponding to the j-th adversarial text sample. Denote the third cosine similarity between the original encoding vector corresponding to the i-th premise text sample and the negative example encoding vector corresponding to the j-th adversarial text sample. Since the value range of j is from 1 to M, M third cosine similarities can be obtained. M represents the number of sample triples included in the text sample set. τ represents the temperature coefficient. For example, it can be taken as 0.05.

[0246] In formula (1), sim(·) represents similarity calculation. Taking cosine similarity as an example, that is, the similarity between two encoding vectors is calculated in the manner provided by formula (2), which will not be elaborated here.

[0247] Based on formula (4), the loss values corresponding to the M sample triples can be summed in the following way, that is:

[0248]

[0249] where L Sup-SimCSE represents the target loss value.

[0250] Based on this, the model parameters of the text encoding model to be trained are optimized using the target loss value until the model training conditions are met, and the text encoding model is obtained.

[0251] Again, in the embodiments of the present application, a method for calculating the loss value based on a supervised loss function is provided. Through the above method, the contrastive learning loss function is used to minimize the distance between positive example pairs and maximize the distance between negative example pairs. Thereby, the model is enabled to learn to generate text vectors with good similarity metrics.

[0252] Optionally, based on one or more of the above Figure 5 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, it may further include:

[0253] Obtain a text sample set, where the text sample set includes N text samples and M sample triples, each sample triple includes a premise text sample, an inheritance text sample, and an adversarial text sample, N is an integer greater than 1, and M is an integer greater than or equal to 1;

[0254] Perform data augmentation processing on each text sample in the text sample set to obtain the positive example encoding vector corresponding to each text sample;

[0255] Obtain N negative example encoding vectors corresponding to each text sample in the text sample set, where the N negative example encoding vectors are generated according to the N text samples;

[0256] Based on the positive example encoding vectors corresponding to each text sample and the N negative example encoding vectors, update the model parameters of the text encoding model to be trained to obtain the target text encoding model;

[0257] For each sample triple, generate the original encoding vector corresponding to the premise text sample, the positive example encoding vector corresponding to the inheritance text sample, and the negative example encoding vector corresponding to the opposing text sample;

[0258] Based on the original encoding vector corresponding to the premise text sample, the positive example encoding vector corresponding to the inheritance text sample, and the negative example encoding vector corresponding to the opposing text sample in each sample triple, fine-tune the target text encoding model to obtain the text encoding model.

[0259] In one or more embodiments, a method for generating a text encoding model based on semi-supervised contrastive learning is introduced. As can be seen from the foregoing embodiments, a set of text samples for training is obtained. The set of text samples includes at least two text samples and at least one sample triple. Among them, the set of text samples can be understood as a batch of data.

[0260] Specifically, for ease of understanding, please refer to Figure 9 , Figure 9 is a schematic diagram of semi-supervised contrastive learning in the embodiments of the present application. As shown in the figure, in the pre-training stage, taking a certain text sample in the set of text samples as an example, assuming that the text sample is "Two little dogs are running". Based on this, perform data augmentation processing on the text sample to generate the positive example encoding vector corresponding to the text sample. At the same time, use the encoding vectors obtained by encoding other text samples as the negative example encoding vectors of this text sample (i.e., "Two little dogs are running"), that is, obtain N negative example encoding vectors. Based on this, calculate the loss value corresponding to each text sample according to the positive example encoding vector corresponding to each text sample and the N negative example encoding vectors, and sum the N loss values to obtain the target loss value. Thus, optimize the model parameters of the text encoding model to be trained by the method of backpropagation and gradient descent to obtain the target text encoding model. The target text encoding model learns a certain degree of text representation.

[0261] Next, enter the fine-tuning stage. Taking a sample triple in the text sample set as an example, assume that the premise text sample is "A person is walking in the park", its corresponding inherited text sample is "A person is walking in the park", and its corresponding adversarial text sample is "A person is swimming in the sea". Based on this, the original encoding vector corresponding to the premise text sample, the positive example encoding vector corresponding to the inherited text sample, and the negative example encoding vector corresponding to the adversarial text sample can be obtained through the target text encoding model. Then, according to the original encoding vector corresponding to the premise text sample, the positive example encoding vector corresponding to the inherited text sample, and the negative example encoding vector corresponding to the adversarial text sample in each sample triple, calculate the loss value corresponding to each sample triple, and sum the M loss values to obtain the target loss value. Thus, the model parameters of the target text encoding model are fine-tuned by the method of backpropagation and gradient descent to obtain the text encoding model. The text encoding model further optimizes the text representation according to the label data.

[0262] Secondly, in the embodiments of the present application, a method for generating a text encoding model based on semi-supervised contrast learning is provided. Through the above method, combining the advantages of unsupervised contrast learning and supervised contrast learning, the model can improve the performance of the model in the case of scarce data.

[0263] Optionally, based on one or more corresponding Figure 5 embodiments, in another optional embodiment provided by the embodiments of the present application, according to K text vectors to be measured and T abnormal text vectors, a detection result for the account to be detected is generated, which may specifically include:

[0264] For each of the T abnormal text vectors, calculate the similarity between the abnormal text vector and each text vector to be measured, and obtain K similarities. The greater the similarity, the higher the similarity degree between the abnormal text vector and the text vector to be measured;

[0265] For each of the T abnormal text vectors, according to the K similarities, determine P text vectors to be measured with the largest similarity to the abnormal text vector, where P is an integer greater than or equal to 1 and less than K;

[0266] According to the P text vectors to be measured corresponding to each abnormal text vector, determine the number of text vectors to be measured containing abnormal content among the K text vectors to be measured;

[0267] Generate a detection result for the account to be detected according to the number of text vectors to be measured containing abnormal content.

[0268] In one or more embodiments, a method for generating a detection result based on vector similarity is introduced. As can be seen from the foregoing embodiments, each abnormal text vector can be used as a reference, and the similarity between each text vector to be tested and each abnormal text vector can be calculated separately. In this application, the cosine similarity is taken as an example for introduction. In practical applications, other metrics can also be used to measure the similarity between text vectors, such as Euclidean distance, cosine distance, etc., which are not limited herein.

[0269] Specifically, taking an abnormal text vector as an example, first, calculate the similarity between this abnormal text vector and K text vectors to be tested respectively, thereby obtaining K similarities. Based on this, select the top P largest similarities from the K similarities, and use the text vectors to be tested corresponding to these P similarities as the P text vectors to be tested that are similar to the abnormal text vector. Taking T as 5, P as 3, and K as 10 as an example, for the convenience of understanding, please refer to Table 1. Table 1 is a schematic diagram of the P text vectors to be tested that are similar to each abnormal text vector.

[0270] Table 1

[0271] Abnormal text vector 1 Text vectors to be measured 1, Text vectors to be measured 2, Text vectors to be measured 6 Abnormal text vector 2 Text vectors to be measured 1, Text vectors to be measured 3, Text vectors to be measured 9 Abnormal text vector 3 Text vectors to be measured 2, Text vectors to be measured 3, Text vectors to be measured 4 Abnormal text vector 4 Text vectors to be measured 1, Text vectors to be measured 4, Text vectors to be measured 8 Abnormal text vector 5 Text vectors to be measured 1, Text vectors to be measured 8, Text vectors to be measured 9

[0272] Among them, text vector to be tested 1, text vector to be tested 2, text vector to be tested 3, text vector to be tested 4, text vector to be tested 6, text vector to be tested 8, and text vector to be tested 9 are relatively similar to the abnormal text vector respectively. Therefore, it can be considered that among the 10 text vectors to be tested, 7 text vectors to be tested may contain abnormal content.

[0273] It should be noted that in practical applications, each text to be tested that may contain abnormal content can be further filtered. For example, match the text to be tested with a blacklist library. If a word in the blacklist library is hit, it means that the text to be tested contains abnormal content. Another example is that the abnormal content of the text to be tested can be determined by means of manual review.

[0274] Exemplarily, if the proportion of the number of text vectors to be tested containing abnormal content among the K text vectors to be tested is less than the first abnormal proportion threshold (for example, 50%), it means that the detection result of the account to be detected is a "normal account".

[0275] Exemplarily, if the proportion of the number of text vectors to be tested containing abnormal content among the K text vectors to be tested is greater than or equal to the first abnormal proportion threshold (for example, 50%) and less than the second abnormal proportion threshold (for example, 70%), it means that the detection result of the account to be detected is a "suspicious account".

[0276] Exemplarily, if the proportion of the number of test texts containing abnormal content among the K test texts is greater than or equal to the second abnormal proportion threshold (e.g., 70%), it indicates that the detection result of the account to be detected is a "malicious account".

[0277] Secondly, in the embodiments of the present application, a method for generating a detection result based on vector similarity is provided. Through the above method, taking each abnormal text vector as a reference, the P test text vectors with the largest similarity are respectively grouped into one category. Thus, it is used as a basis for identifying whether the account is abnormal, thereby improving the feasibility and operability of the solution.

[0278] Optionally, based on one or more of the above Figure 5 In another optional embodiment provided by the embodiments of the present application, based on the K test text vectors and the T abnormal text vectors, a detection result for the account to be detected is generated, which may specifically include:

[0279] Calculate the similarity between each of the K test text vectors and each of the T abnormal text vectors respectively to obtain K×T similarities. Among them, the greater the similarity, the higher the degree of similarity between the abnormal text vector and the test text vector;

[0280] If there is at least one similarity greater than or equal to the similarity threshold among the K×T similarities, determine the number of test texts containing abnormal content among the K test texts according to the at least one similarity;

[0281] Generate a detection result for the account to be detected according to the number of test texts containing abnormal content.

[0282] In one or more embodiments, a method for generating a detection result based on vector similarity is introduced. As can be seen from the foregoing embodiments, each abnormal text vector can be used as a reference to calculate the similarity between each test text vector and each abnormal text vector respectively. In the present application, the cosine similarity is taken as an example for introduction. In practical applications, other metrics can also be used to measure the similarity between text vectors, which is not limited herein.

[0283] Specifically, first, calculate the similarity between each abnormal text vector and each test text vector respectively, thereby obtaining K×T similarities. Then, obtain at least one similarity greater than or equal to the similarity threshold (e.g., 0.8) from the K×T similarities. Taking T as 5 and K as 10 as an example, for ease of understanding, please refer to Table 2. Table 2 is an illustration of the test text vectors with a similarity greater than or equal to the similarity threshold to each abnormal text vector.

[0284] Table 2

[0285] Abnormal text vector 1 Text vectors to be measured 1, Text vectors to be measured 2 Abnormal text vector 2 Text vector to be measured 1 Abnormal text vector 3 Text vector to be measured 4 Abnormal text vector 4 None Abnormal text vector 5 Text vectors to be measured 1, Text vectors to be measured 8, Text vectors to be measured 9

[0286] Among them, the similarity between the text vector to be measured 1, the text vector to be measured 2, the text vector to be measured 4, the text vector to be measured 8, and the text vector to be measured 9 and the abnormal text vector is greater than or equal to the similarity threshold. Therefore, it can be considered that among the 10 text vectors to be measured, 5 text vectors to be measured corresponding to the text to be measured contain abnormal content. That is, the proportion of the number of text vectors to be measured containing abnormal content among the K text vectors to be measured is 50%.

[0287] Exemplarily, if the proportion of the number of text vectors to be measured containing abnormal content among the K text vectors to be measured is less than the first abnormal proportion threshold (for example, 50%), it indicates that the detection result of the account to be detected is a "normal account".

[0288] Exemplarily, if the proportion of the number of text vectors to be measured containing abnormal content among the K text vectors to be measured is greater than or equal to the first abnormal proportion threshold (for example, 50%) and less than the second abnormal proportion threshold (for example, 70%), it indicates that the detection result of the account to be detected is a "suspicious account".

[0289] Exemplarily, if the proportion of the number of text vectors to be measured containing abnormal content among the K text vectors to be measured is greater than or equal to the second abnormal proportion threshold (for example, 70%), it indicates that the detection result of the account to be detected is a "malicious account".

[0290] Secondly, in the embodiments of the present application, a method for generating a detection result based on vector similarity is provided. Through the above method, taking each abnormal text vector as a standard, the text vectors to be measured with a sufficiently large similarity are classified into one category respectively. Thus, it is used as a basis for identifying whether the account is abnormal, thereby improving the feasibility and operability of the solution.

[0291] Optionally, on the basis of one or more corresponding embodiments above, in another optional embodiment provided by the embodiments of the present application, according to the K text vectors to be measured and the T abnormal text vectors, a detection result for the account to be detected is generated, which may specifically include: Figure 5 Initialize Q clustering centers, where Q is an integer greater than 1;

[0292] Calculate the Euclidean distance between each text vector to be measured among the K text vectors to be measured and each abnormal text vector among the T abnormal text vectors and each clustering center respectively, and obtain (K + T) × Q Euclidean distances;

[0293] According to the (K + T) × Q Euclidean distances, divide each text vector to be measured and each abnormal text vector into the clustering clusters corresponding to the clustering centers with the smallest Euclidean distance respectively, and obtain Q clustering clusters;

[0294]

[0295] According to Q clusters, each text vector to be measured and each abnormal text vector are divided until the clustering optimization condition is satisfied, and Q target clusters are obtained;

[0296] Generate a detection result for the account to be detected according to the number of abnormal text vectors included in each target cluster among the Q target clusters.

[0297] In one or more embodiments, a method for generating an account detection result based on the K-Means clustering algorithm is introduced. As can be seen from the foregoing embodiments, after obtaining K text vectors to be measured and T abnormal text vectors, similar text vectors can be grouped through the K-Means clustering algorithm.

[0298] Specifically, the K text vectors to be measured and the T abnormal text vectors are used as the text vector set to be clustered, that is, X = {X 1 , X 2 , X 3 ,..., X K+T}, where each text vector (that is, the text vector to be measured, the abnormal text vector) has m dimensions (for example, 128 dimensions). The goal of the K-Means clustering algorithm is to cluster into the specified Q clusters according to the similarity between the K text vectors to be measured and the T abnormal text vectors, and each text vector belongs to only one cluster that is closest to the cluster center. First, initialize Q cluster centers, that is, {C 1 , C 2 , C 3 ,..., C Q}, where 1 < Q ≤ K + T. Then, the following method can be used to calculate the Euclidean distance between each text vector (that is, the text vector to be measured and the abnormal text vector) in the text vector set X and each cluster center:

[0299]

[0300] Among them, X i represents the i-th text vector in the text vector set X (that is, the text vector to be measured, the abnormal text vector), and 1 ≤ i ≤ K + T. C j represents the j-th cluster center, and 1 ≤ j ≤ Q. t represents the t-th dimension. m represents the number of dimensions of the text vector. X it represents the value of the t-th dimension of the i-th text vector (that is, the text vector to be measured, the abnormal text vector). C jt represents the value of the t-th dimension of the j-th cluster center.

[0301] Combined with formula (6), (K + T) × Q Euclidean distances can be calculated. Based on this, the distances from each text vector (i.e., the text vector to be measured, the abnormal text vector) to each cluster center are compared in sequence, and each text vector is respectively assigned to the cluster corresponding to the cluster center with the smallest Euclidean distance, obtaining Q clusters, that is, {S 1 , S 2 , S 3 ,..., S Q}.

[0302] Next, update the cluster center of each of the Q clusters. The updated cluster center is the mean value of all text vectors in the cluster in each dimension. The cluster center can be calculated in the following way:

[0303]

[0304] where C v represents the v-th cluster center, and 1 ≤ v ≤ K + T. |S v | represents the number of text vectors in the v-th cluster. X i represents the i-th text vector in the v-th cluster (i.e., the text vector to be measured, the abnormal text vector), and 1 ≤ i ≤ |S v |.

[0305] After updating the cluster center of each cluster, calculate the Euclidean distance between each text vector (i.e., the text vector to be measured, the abnormal text vector) in the text vector set and each cluster center again. Then, assign each text vector to the cluster corresponding to the cluster center with the smallest Euclidean distance. Until the clustering optimization condition is met, Q target clusters are obtained. It should be noted that the clustering optimization condition can be that the cluster center no longer changes, or the sum of squared errors is locally minimized.

[0306] After obtaining the Q target clusters, count the number of abnormal text vectors included in each target cluster. Taking Q = 3 as an example, for the sake of understanding, please refer to Table 3. Table 3 is an example of the text vectors included in each target cluster.

[0307] Table 3

[0308] Target cluster identification Number of text vectors to be measured Number of abnormal text vectors Proportion of abnormal text vectors A 5 95 95% B 20 180 90% C 20 30 60%

[0309] If the proportion of abnormal text vectors in the target clustering cluster is greater than or equal to the in-cluster proportion threshold (e.g., 90%), it indicates that the target clustering cluster belongs to an abnormal clustering cluster. Based on this, the detection result of the account to be detected can be generated according to the proportion of the text vectors to be tested belonging to the abnormal clustering cluster among the K text vectors to be tested. Taking Table 3 as an example, K is equal to 45 (i.e., 5 + 20 + 20 = 45), where 25 text vectors to be tested belong to the abnormal clustering cluster (i.e., the target clustering cluster A and the target clustering cluster B). Therefore, the proportion of the text vectors to be tested belonging to the abnormal clustering cluster is 55.6% (i.e., 25 ÷ 45).

[0310] Exemplarily, if the proportion of the text vectors to be tested belonging to the abnormal clustering cluster is less than the first proportion threshold (e.g., 50%), it indicates that the detection result of the account to be detected is a "normal account".

[0311] Exemplarily, if the proportion of the text vectors to be tested belonging to the abnormal clustering cluster is greater than or equal to the first proportion threshold (e.g., 50%) and less than the second proportion threshold (e.g., 70%), it indicates that the detection result of the account to be detected is a "suspicious account".

[0312] Exemplarily, if the proportion of the text vectors to be tested belonging to the abnormal clustering cluster is greater than or equal to the second proportion threshold (e.g., 70%), it indicates that the detection result of the account to be detected is a "malicious account".

[0313] Furthermore, the sensitive content involved in the target clustering cluster A and the sensitive content involved in the target clustering cluster B can be separately marked, thereby determining the portrait data of the target clustering cluster A and the portrait data of the target clustering cluster B. Further, the type of the account to be detected can be determined.

[0314] Secondly, in the embodiments of the present application, a method for generating an account detection result based on the K-Means clustering algorithm is provided. Through the above method, the K-Means clustering algorithm can overcome the inaccuracy of clustering with a small number of samples, and perform iterative correction on the already obtained clustering to determine the clustering of some samples, thereby optimizing the problem of unreasonable classification of the initial supervised learning samples. In addition, since only some small samples are clustered, the overall clustering time complexity can be reduced.

[0315] Optionally, based on one or more corresponding embodiments above Figure 5 In another optional embodiment provided by the embodiments of the present application, according to the K text vectors to be tested and the T abnormal text vectors, the detection result for the account to be detected is generated, which may specifically include:

[0316] Regarding the K text vectors to be tested and the T abnormal text vectors as a text vector set, where the text vector set includes (K + T) text vectors;

[0317] Select a text vector from the set of text vectors as the first text vector;

[0318] Calculate the similarity between the first text vector and the remaining (K + T - 1) text vectors in the set of text vectors respectively, and obtain the second text vector corresponding to the maximum similarity;

[0319] If the maximum similarity is greater than or equal to the target threshold, add the second text vector to the first cluster corresponding to the first text vector, and update the cluster center of the first cluster;

[0320] If the maximum similarity is less than the target threshold, use the second text vector as the cluster center of the second cluster;

[0321] When each text vector in the set of text vectors is divided into the corresponding cluster, obtain R clusters, where R is an integer greater than or equal to 1;

[0322] Generate a detection result for the account to be detected according to the number of abnormal text vectors included in each of the R clusters.

[0323] In one or more embodiments, a method for generating an account detection result based on the single pass clustering algorithm is provided. As can be seen from the foregoing embodiments, after obtaining K text vectors to be measured and T abnormal text vectors, similar text vectors can be grouped through the single pass clustering algorithm.

[0324] Specifically, take the K text vectors to be measured and the T abnormal text vectors as the set of text vectors to be clustered, that is, X = {X 1 , X 2 , X 3 ,..., X K+T}, where each text vector (that is, the text vector to be measured, the abnormal text vector) has m dimensions (for example, 128 dimensions). First, randomly select a text vector from the set of text vectors as the first text vector. Then, calculate the similarity between the first text vector and each of the remaining text vectors in the set of text vectors, and take the text vector with the maximum similarity as the second text vector. If the maximum similarity is greater than or equal to the target threshold, add the second text vector to the first cluster corresponding to the first text vector. If the maximum similarity is less than the target threshold, create a new cluster (that is, the second cluster), and classify the second text vector into the second cluster.

[0325] Assume that the set of text vectors includes 5 text vectors (that is, the text vector to be measured, the abnormal text vector). For ease of understanding, please refer to Table 4, and Table 4 is a schematic of each text vector in the set of text vectors.

[0326] Table 4

[0327] Text vector 1 (1,3,3,2,2) Text vector 2 (2,1,0,1,2) Text vector 3 (0,2,0,0,1) Text vector 4 (0,3,0,3,5) Text vector 5 (1,0,1,0,1)

[0328] Taking the text vector 1 as the first text vector as an example, assume that its cluster center is C1. At this time, the first cluster only contains the text vector 1. Therefore, the text vector 1 can be represented as the cluster center C1, that is, C1 = (1, 3, 3, 2, 2). Then, select the next text vector (i.e., text vector 2) as the second text vector, and calculate the similarity between the text vector 2 and the current cluster centers. Since there is only the cluster center C1 at this time, the similarity between the text vector 2 and the cluster center C1 can be calculated. Taking the dot product as the similarity calculation method, that is:

[0329] sim(Doc2, C1) = 1 * 2 + 1 * 3 + 0 * 3 + 1 * 2 + 2 * 2 = 11;

[0330] where Doc2 represents the text vector 2.

[0331] If the target threshold is 10, then sim(Doc2, C1) is greater than the target threshold. Therefore, add the text vector 2 to the first cluster corresponding to the cluster center C1. Based on this, it is necessary to recalculate the cluster center of the first cluster. Since the first cluster includes the text vector 1 and the text vector 2 at this time, the cluster center can be updated by using the average method, that is:

[0332] C’1 = ((1 + 2) / 2, (3 + 1) / 2, (3 + 0) / 2, (2 + 1) / 2, (2 + 2) / 2) = (1.5, 2, 1.5, 1.5, 2);

[0333] Next, select the text vector 3 for calculation. Calculate the similarity between the text vector 3 and the current cluster centers. Since there is only the cluster center C1 at this time, the similarity between the text vector 3 and the updated cluster center C’1 can be calculated. Taking the dot product as the similarity calculation method, that is:

[0334] sim(Doc3, C’1) = 0 * 1.5 + 2 * 2 + 0 * 1.5 + 0 * 1.5 + 1 * 2 = 6;

[0335] where Doc3 represents the text vector 3.

[0336] It can be seen that sim(Doc3, C’1) is less than the target threshold. Therefore, use the text vector 3 to generate a new cluster. And so on, perform the above processing on each text vector in the text vector set until each text vector belongs to the corresponding cluster.

[0337] Taking the example of finally obtaining R clustering clusters, count the number of abnormal text vectors included in each clustering cluster respectively. If the proportion of abnormal text vectors in the clustering cluster is greater than or equal to the in-cluster proportion threshold (for example, 90%), it indicates that the clustering cluster belongs to an abnormal clustering cluster. Based on this, the detection result of the account to be detected can be generated according to the proportion of the text vectors to be detected that belong to the abnormal clustering clusters among the K text vectors to be detected. The specific method can refer to the foregoing embodiments and will not be elaborated here.

[0338] Secondly, in the embodiments of the present application, a method for generating an account detection result based on a single-pass clustering algorithm is provided. Through the above method, the single pass clustering algorithm does not need to specify the number of clusters, can achieve incremental and dynamic clustering of streaming data, is suitable for mining streaming data, and the efficiency of the algorithm is relatively high. Thus, the efficiency of account detection can be improved.

[0339] Optionally, on the basis of the foregoing Figure 5 corresponding one or more embodiments, in another optional embodiment provided by the embodiments of the present application, after generating the detection result for the account to be detected according to the K text vectors to be detected and the T abnormal text vectors, it may further include:

[0340] In the case where the detection result indicates that the account to be detected is a suspicious account, intercept the text sent by the suspicious account;

[0341] In the case where the detection result indicates that the account to be detected is a malicious account, intercept the text sent by the malicious account and ban the usage permission of the malicious account.

[0342] In one or more embodiments, a method for performing gradient processing on an account based on the detection result is introduced. As can be seen from the foregoing embodiments, based on the K text vectors to be detected and the T abnormal text vectors, similarity calculation or clustering can be performed, and then the detection result of the account to be detected can be determined. The following will take the account to be detected being a suspicious account and a malicious account as examples for introduction.

[0343] I. The account to be detected is a suspicious account;

[0344] Exemplarily, for the sake of easy understanding, please refer to Figure 10 , Figure 10 which is a schematic diagram for processing a suspicious account in the embodiments of the present application. As shown in Figure (A) of Figure 10 , user B inputs the greeting text indicated by D1 (for example, "What a coincidence. We seem to be junior high school classmates. Add me."). Then, click the send control indicated by D2. Thus, user B initiates a friend addition request to user A. If the account to be detected (i.e., user B's account) is a suspicious account, it will be displayed as Figure 10The interface shown in Figure (B). Among them, D3 is used to indicate the first prompt message, and the first prompt message indicates that the greeting text input by user B has been intercepted, that is, user A will not receive this greeting text.

[0345] Second, the account to be detected is a malicious account;

[0346] Exemplarily, for ease of understanding, please refer to Figure 11 , Figure 11 which is a schematic diagram for processing a malicious account in an embodiment of the present application. As shown in Figure (A), user B inputs the greeting text indicated by E1 (for example, "What a coincidence. We seem to be junior high school classmates. Add me."). Then, click the send control indicated by E2. Thus, user B sends a friend addition application to user A. If the account to be detected (that is, user B's account) is a malicious account, the interface shown in Figure (B) will be displayed. Among them, E3 is used to indicate the second prompt message, and the second prompt message indicates that user B's account has been blocked and the greeting text input by user B has been intercepted, that is, user A will not receive this greeting text. Figure 11 Figure 11

[0347]

[0348] Secondly, in the embodiment of the present application, a method for gradient processing of an account based on the detection result is provided. Through the above method, corresponding processing modes are adopted according to the abnormal degree of the account. On the one hand, it can avoid overly severe processing of normal accounts. On the other hand, it can timely process malicious accounts and improve the security of social interaction. Figure 12 Figure 12 The business processing device in the present application will be described in detail below. Please refer to Figure 12 , Figure 12 which is a schematic diagram of an embodiment of the business processing device in an embodiment of the present application. The business processing device 30 includes:

[0349] An acquisition module 310, configured to acquire a set of texts to be tested, where the set of texts to be tested includes K texts to be tested sent by the account to be detected, and K is an integer greater than or equal to 1;

[0350] An encoding module 320, configured to encode each text to be tested in the set of texts to be tested by using a text encoding model to obtain K text vectors to be tested, where the K text vectors to be tested have a one-to-one correspondence with the K texts to be tested, and the text encoding model is obtained by contrastive learning using positive example encoding vector pairs and negative example encoding vector pairs;

[0351] The acquisition module 310 is further configured to acquire a set of abnormal texts, where the set of abnormal texts includes T abnormal texts that have been marked as abnormal content, and T is an integer greater than or equal to 1;

[0352] The encoding module 320 is further configured to encode each abnormal text in the abnormal text set by using a text encoding model to obtain T abnormal text vectors, where the T abnormal text vectors have a one-to-one correspondence with the T abnormal texts;

[0353] The generation module 330 is configured to generate a detection result for the account to be detected according to the K text vectors to be measured and the T abnormal text vectors, where the detection result is used to represent the abnormal degree of the account to be detected.

[0354] Optionally, based on the above Figure 12 In another embodiment of the service processing device 30 provided in the embodiment of the present application, on the basis of the corresponding embodiment, the service processing device 30 further includes a processing module 340 and a training module 350;

[0355] The acquisition module 310 is further configured to acquire a text sample set, where the text sample set includes N text samples, and N is an integer greater than 1;

[0356] The processing module 340 is configured to perform data enhancement processing on each text sample in the text sample set to obtain a positive example encoding vector corresponding to each text sample;

[0357] The acquisition module 310 is further configured to acquire N negative example encoding vectors corresponding to each text sample in the text sample set, where the N negative example encoding vectors are generated according to the N text samples;

[0358] The processing module 340 is further configured to calculate a loss value corresponding to each text sample according to the positive example encoding vector and the N negative example encoding vectors corresponding to the text sample for each text sample;

[0359] The training module 350 is configured to update the model parameters of the text encoding model to be trained according to the loss value corresponding to each text sample in the text sample set until the model training condition is met, and obtain the text encoding model.

[0360] Optionally, based on the above Figure 12 In another embodiment of the service processing device 30 provided in the embodiment of the present application, on the basis of the corresponding embodiment,

[0361] Specifically, for each text sample, the processing module 340 encodes the text sample twice through the text encoding model to be trained to obtain an original encoding vector and a positive example encoding vector corresponding to the text sample.

[0362] Optionally, based on the above Figure 12 In another embodiment of the service processing device 30 provided in the embodiment of the present application, on the basis of the corresponding embodiment,

[0363] The processing module 340 is specifically configured to perform a synonym replacement process on at least one text unit in the text sample for each text sample, so as to obtain the target text sample corresponding to the text sample;

[0364] For each text sample, encode the text sample through the text encoding model to be trained, so as to obtain the original encoding vector corresponding to the text sample;

[0365] For each text sample, encode the target text sample corresponding to the text sample through the text encoding model to be trained, so as to obtain the positive example encoding vector corresponding to the text sample.

[0366] Optionally, based on the corresponding embodiment above, Figure 12 In another embodiment of the service processing device 30 provided by the embodiment of the present application,

[0367] The processing module 340 is specifically configured to perform an addition and deletion process on the text sample for each text sample, so as to obtain the target text sample corresponding to the text sample, where the addition and deletion process includes deleting at least one text unit in the text sample, or adding at least one text unit to the text sample;

[0368] For each text sample, encode the text sample through the text encoding model to be trained, so as to obtain the original encoding vector corresponding to the text sample;

[0369] For each text sample, encode the target text sample corresponding to the text sample through the text encoding model to be trained, so as to obtain the positive example encoding vector corresponding to the text sample.

[0370] Optionally, based on the corresponding embodiment above, Figure 12 In another embodiment of the service processing device 30 provided by the embodiment of the present application,

[0371] The processing module 340 is specifically configured to determine the cosine similarity between the original encoding vector corresponding to the text sample and the positive example encoding vector for each text sample;

[0372] For each text sample, determine N cosine similarities between the original encoding vector corresponding to the text sample and N negative example encoding vectors;

[0373] For each text sample, calculate the loss value corresponding to the text sample according to the cosine similarity between the original encoding vector corresponding to the text sample and the positive example encoding vector, and the N cosine similarities between the original encoding vector corresponding to the text sample and the N negative example encoding vectors.

[0374] Optionally, based on the corresponding embodiment above, Figure 12Based on the corresponding embodiment, in another embodiment of the service processing device 30 provided by the embodiments of the present application,

[0375] The acquisition module 310 is further configured to acquire a text sample set, where the text sample set includes M sample triples, and each sample triple includes a premise text sample, an inheritance text sample, and an opposing text sample, and M is an integer greater than or equal to 1;

[0376] The generation module 330 is further configured to generate, for each sample triple, an original encoding vector corresponding to the premise text sample, a positive example encoding vector corresponding to the inheritance text sample, and a negative example encoding vector corresponding to the opposing text sample;

[0377] The processing module 340 is further configured to calculate, for each sample triple, a loss value corresponding to the sample triple according to the original encoding vector corresponding to the premise text sample, the positive example encoding vector corresponding to the inheritance text sample, and the negative example encoding vector corresponding to the opposing text sample;

[0378] The training module 350 is further configured to update the model parameters of the text encoding model to be trained according to the loss value corresponding to each sample triple in the text sample set until the model training condition is satisfied, and obtain the text encoding model.

[0379] Optionally, based on the above Figure 12 Based on the corresponding embodiment, in another embodiment of the service processing device 30 provided by the embodiments of the present application,

[0380] The generation module 330 is specifically configured to, for each sample triple, encode the premise text sample in the sample triple through the text encoding model to be trained to obtain an original encoding vector corresponding to the premise text sample;

[0381] For each sample triple, encode the inheritance text sample in the sample triple through the text encoding model to be trained to obtain a positive example encoding vector corresponding to the inheritance text sample;

[0382] For each sample triple, encode the opposing text sample in the sample triple through the text encoding model to be trained to obtain a negative example encoding vector corresponding to the opposing text sample.

[0383] Optionally, based on the above Figure 12 Based on the corresponding embodiment, in another embodiment of the service processing device 30 provided by the embodiments of the present application,

[0384] The generation module 330 is specifically configured to, for each sample triple, encode each premise text sample in the sample triple at least twice through the text encoding model to be trained, so as to obtain at least two original encoding vectors corresponding to the premise text sample;

[0385] For each sample triple, encode each inheritance text sample in the sample triple at least twice through the text encoding model to be trained, so as to obtain at least two positive example encoding vectors corresponding to the inheritance text sample;

[0386] For each sample triple, encode each opposing text sample in the sample triple at least twice through the text encoding model to be trained, so as to obtain at least two negative example encoding vectors corresponding to the opposing text sample.

[0387] Optionally, based on the foregoing Figure 12 corresponding embodiment, in another embodiment of the service processing device 30 provided by the embodiments of the present application,

[0388] The processing module 340 is specifically configured to, for each sample triple, determine the first cosine similarity between the original encoding vector corresponding to the premise text sample and the positive example encoding vector corresponding to the inheritance text sample;

[0389] For each sample triple, determine M second cosine similarities between the original encoding vector corresponding to the premise text sample and the positive example encoding vectors corresponding to M inheritance text samples;

[0390] For each sample triple, determine M third cosine similarities between the original encoding vector corresponding to the premise text sample and the negative example encoding vectors corresponding to M opposing text samples;

[0391] For each sample triple, calculate the loss value corresponding to the sample triple according to the first cosine similarity, the M second cosine similarities, and the M third cosine similarities.

[0392] Optionally, based on the foregoing Figure 12 corresponding embodiment, in another embodiment of the service processing device 30 provided by the embodiments of the present application,

[0393] The acquisition module 310 is further configured to acquire a text sample set, where the text sample set includes N text samples and M sample triples, each sample triple includes a premise text sample, an inheritance text sample, and an opposing text sample, N is an integer greater than 1, and M is an integer greater than or equal to 1;

[0394] The processing module 340 is further configured to perform data augmentation processing on each text sample in the text sample set to obtain a positive example encoding vector corresponding to each text sample;

[0395] The obtaining module 310 is further configured to obtain N negative example encoding vectors corresponding to each text sample in the text sample set, where the N negative example encoding vectors are generated according to N text samples;

[0396] The training module 350 is further configured to update the model parameters of the text encoding model to be trained according to the positive example encoding vector and the N negative example encoding vectors corresponding to each text sample, so as to obtain the target text encoding model;

[0397] The generating module 330 is further configured to generate, for each sample triple, an original encoding vector corresponding to the premise text sample, a positive example encoding vector corresponding to the inheritance text sample, and a negative example encoding vector corresponding to the opposing text sample;

[0398] The training module 350 is further configured to fine-tune the target text encoding model according to the original encoding vector corresponding to the premise text sample, the positive example encoding vector corresponding to the inheritance text sample, and the negative example encoding vector corresponding to the opposing text sample in each sample triple, so as to obtain the text encoding model.

[0399] Optionally, based on the corresponding embodiment above Figure 12 In another embodiment of the service processing device 30 provided by the embodiment of the present application,

[0400] The generating module 330 is specifically configured to calculate the similarity between each abnormal text vector in the T abnormal text vectors and each text vector to be measured, so as to obtain K similarities, where the greater the similarity, the higher the similarity degree between the abnormal text vector and the text vector to be measured;

[0401] For each abnormal text vector in the T abnormal text vectors, determine P text vectors to be measured with the greatest similarity to the abnormal text vector according to the K similarities, where P is an integer greater than or equal to 1 and less than K;

[0402] Determine the number of text vectors to be measured containing abnormal content among the K text vectors to be measured according to the P text vectors to be measured corresponding to each abnormal text vector;

[0403] Generate a detection result for the account to be detected according to the number of text vectors to be measured containing abnormal content.

[0404] Optionally, based on the corresponding embodiment above Figure 12 In another embodiment of the service processing device 30 provided by the embodiment of the present application,

[0405] The generation module 330 is specifically configured to calculate the similarity between each of the K text vectors to be measured and each of the T abnormal text vectors respectively, obtaining K×T similarities. Among them, the greater the similarity, the higher the similarity degree between the abnormal text vector and the text vector to be measured;

[0406] If there is at least one similarity greater than or equal to the similarity threshold among the K×T similarities, then determine the number of text vectors to be measured containing abnormal content among the K text vectors to be measured according to the at least one similarity;

[0407] Generate a detection result for the account to be detected according to the number of text vectors to be measured containing abnormal content.

[0408] Optionally, on the basis of the above Figure 12 corresponding embodiment, in another embodiment of the service processing device 30 provided by the embodiment of the present application,

[0409] The generation module 330 is specifically configured to initialize Q clustering centers, where Q is an integer greater than 1;

[0410] Calculate the Euclidean distance between each of the K text vectors to be measured and each of the T abnormal text vectors and each clustering center respectively, obtaining (K + T)×Q Euclidean distances;

[0411] According to the (K + T)×Q Euclidean distances, divide each text vector to be measured and each abnormal text vector into the clustering cluster corresponding to the clustering center with the smallest Euclidean distance respectively, obtaining Q clustering clusters;

[0412] According to the Q clustering clusters, divide each text vector to be measured and each abnormal text vector until the clustering optimization condition is satisfied, obtaining Q target clustering clusters;

[0413] Generate a detection result for the account to be detected according to the number of abnormal text vectors included in each target clustering cluster among the Q target clustering clusters.

[0414] Optionally, on the basis of the above Figure 12 corresponding embodiment, in another embodiment of the service processing device 30 provided by the embodiment of the present application,

[0415] The generation module 330 is specifically configured to use the K text vectors to be measured and the T abnormal text vectors as a text vector set, where the text vector set includes (K + T) text vectors;

[0416] Select a text vector from the text vector set as the first text vector;

[0417] Calculate the similarity between the first text vector and the remaining (K + T - 1) text vectors in the text vector set respectively, and obtain the second text vector corresponding to the maximum similarity;

[0418] If the maximum similarity is greater than or equal to the target threshold, add the second text vector to the first cluster corresponding to the first text vector, and update the cluster center of the first cluster;

[0419] If the maximum similarity is less than the target threshold, use the second text vector as the cluster center of the second cluster;

[0420] When each text vector in the text vector set is divided into the corresponding cluster, R clusters are obtained, where R is an integer greater than or equal to 1;

[0421] Generate a detection result for the account to be detected according to the number of abnormal text vectors included in each of the R clusters.

[0422] Optionally, based on the above Figure 12 In another embodiment of the service processing device 30 provided by the embodiment of the present application, on the basis of the corresponding embodiment,

[0423] The processing module 340 is further configured to, after generating a detection result for the account to be detected according to the K text vectors to be measured and the T abnormal text vectors, perform an interception process on the text sent by the suspicious account when the detection result indicates that the account to be detected is a suspicious account;

[0424] The processing module 340 is further configured to, when the detection result indicates that the account to be detected is a malicious account, perform an interception process on the text sent by the malicious account and block the usage permission of the malicious account.

[0425] Figure 13 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 400 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 422 (for example, one or more processors) and a memory 432, and one or more storage media 430 (for example, one or more mass storage devices) for storing application programs 442 or data 444. Among them, the memory 432 and the storage media 430 can be transient storage or persistent storage. The program stored in the storage media 430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the computer device. Further, the central processing unit 422 may be configured to communicate with the storage media 430 and execute a series of instruction operations in the storage media 430 on the computer device 400.

[0426] The computer device 400 may also include one or more power supplies 426, one or more wired or wireless network interfaces 450, one or more input / output interfaces 458, and / or one or more operating systems 441, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.

[0427] The steps performed by the computer device in the above embodiments may be based on the Figure 13 structure of the computer device shown.

[0428] In an embodiment of the present application, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the methods described in the foregoing embodiments are implemented.

[0429] In an embodiment of the present application, a computer program product is further provided, including a computer program. When the computer program is executed by a processor, the steps of the methods described in the foregoing embodiments are implemented.

[0430] It can be understood that in the specific implementation of the present application, when it comes to data related to user information, text to be measured, etc., when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

[0431] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0432] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the devices or units may be in electrical, mechanical, or other forms.

[0433] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0434] In addition, each functional unit in various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0435] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a server, a terminal device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store computer programs.

[0436] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of the present application.

Claims

1. A business processing method, characterized in that: include: Obtain a text set to be tested, wherein the text set to be tested includes K texts to be tested sent through the account to be tested, and K is an integer greater than or equal to 1; Using a text encoding model to encode each test text in the test text set to obtain K test text vectors, wherein the K test text vectors have a one-to-one correspondence with the K test texts, and the text encoding model is obtained by comparative learning using a positive example encoding vector pair and a negative example encoding vector pair; Acquire an abnormal text set, wherein the abnormal text set includes T abnormal texts marked as abnormal contents, where T is an integer greater than or equal to 1; Using the text encoding model to encode each abnormal text in the abnormal text set to obtain T abnormal text vectors, wherein the T abnormal text vectors have a one-to-one correspondence with the T abnormal texts; A detection result for the account to be detected is generated based on the K text vectors to be detected and the T abnormal text vectors, wherein the detection result is used to indicate the abnormality degree of the account to be detected.

2. The method according to claim 1, characterized in that The method further comprises: Acquire a text sample set, wherein the text sample set includes N text samples, and N is an integer greater than 1; Performing data enhancement processing on each text sample in the text sample set to obtain a positive example encoding vector corresponding to each text sample; Obtaining N negative example encoding vectors corresponding to each text sample in the text sample set, wherein the N negative example encoding vectors are generated according to the N text samples; For each text sample, the loss value corresponding to the text sample is calculated according to the positive example encoding vector and N negative example encoding vectors corresponding to the text sample; According to the loss value corresponding to each text sample in the text sample set, the model parameters of the text encoding model to be trained are updated until the model training conditions are met to obtain the text encoding model.

3. The method according to claim 2, characterized in that The performing data enhancement processing on each text sample in the text sample set to obtain a positive example encoding vector corresponding to each text sample includes: For each of the text samples, the text sample is encoded twice using the text encoding model to be trained to obtain an original encoding vector and a positive example encoding vector corresponding to the text sample.

4. The method according to claim 2, characterized in that: The performing data enhancement processing on each text sample in the text sample set to obtain a positive example encoding vector corresponding to each text sample includes: For each of the text samples, performing synonymous replacement processing on at least one text unit in the text sample to obtain a target text sample corresponding to the text sample; For each of the text samples, encoding the text sample by using the text encoding model to be trained to obtain an original encoding vector corresponding to the text sample; For each of the text samples, the target text sample corresponding to the text sample is encoded by the text encoding model to be trained to obtain a positive example encoding vector corresponding to the text sample.

5. The method according to claim 2, characterized in that: The performing data enhancement processing on each text sample in the text sample set to obtain a positive example encoding vector corresponding to each text sample includes: For each of the text samples, adding and deleting the text sample to obtain a target text sample corresponding to the text sample, wherein the adding and deleting processing includes deleting at least one text unit in the text sample, or adding at least one text unit in the text sample; For each of the text samples, encoding the text sample by using the text encoding model to be trained to obtain an original encoding vector corresponding to the text sample; For each of the text samples, the target text sample corresponding to the text sample is encoded by the text encoding model to be trained to obtain a positive example encoding vector corresponding to the text sample.

6. The method according to any one of claims 2 to 5, characterized in that For each of the text samples, according to the positive example encoding vector and N negative example encoding vectors corresponding to the text sample, a loss value corresponding to the text sample is calculated, including: For each text sample, determining the cosine similarity between the original encoding vector corresponding to the text sample and the positive example encoding vector; For each of the text samples, determining N cosine similarities between the original encoding vector corresponding to the text sample and the N negative example encoding vectors; For each text sample, the loss value corresponding to the text sample is calculated based on the cosine similarity between the original encoding vector corresponding to the text sample and the positive example encoding vector, and the N cosine similarities between the original encoding vector corresponding to the text sample and N negative example encoding vectors.

7. The method according to claim 1, characterized in that The method further comprises: Acquire a text sample set, wherein the text sample set includes M sample triples, each sample triple includes a premise text sample, a successor text sample, and an opposing text sample, and M is an integer greater than or equal to 1; For each sample triple, generate an original encoding vector corresponding to the premise text sample, a positive example encoding vector corresponding to the inherited text sample, and a negative example encoding vector corresponding to the opposing text sample; For each sample triple, the loss value corresponding to the sample triple is calculated according to the original encoding vector corresponding to the premise text sample, the positive example encoding vector corresponding to the inherited text sample, and the negative example encoding vector corresponding to the opposing text sample; According to the loss value corresponding to each sample triplet in the text sample set, the model parameters of the text encoding model to be trained are updated until the model training conditions are met to obtain the text encoding model.

8. The method according to claim 7, characterized in that The step of generating, for each sample triplet, an original encoding vector corresponding to the premise text sample, a positive example encoding vector corresponding to the inherited text sample, and a negative example encoding vector corresponding to the opposing text sample comprises: For each sample triple, encoding the premise text sample in the sample triple by the to-be-trained text encoding model to obtain an original encoding vector corresponding to the premise text sample; For each sample triple, encoding the inherited text sample in the sample triple by the to-be-trained text encoding model to obtain a positive example encoding vector corresponding to the inherited text sample; For each sample triplet, the opposing text sample in the sample triplet is encoded by the to-be-trained text encoding model to obtain a negative example encoding vector corresponding to the opposing text sample.

9. The method according to claim 7, characterized in that: The step of generating, for each sample triplet, an original encoding vector corresponding to the premise text sample, a positive example encoding vector corresponding to the inherited text sample, and a negative example encoding vector corresponding to the opposing text sample comprises: For each of the sample triples, encoding the premise text samples in the sample triples at least twice by using the text encoding model to be trained to obtain at least two original encoding vectors corresponding to the premise text samples; For each of the sample triples, encoding the inherited text samples in the sample triples at least twice by using the to-be-trained text encoding model to obtain at least two positive example encoding vectors corresponding to the inherited text samples; For each sample triplet, the opposing text samples in the sample triplet are respectively encoded at least twice by the to-be-trained text encoding model to obtain at least two negative example encoding vectors corresponding to the opposing text samples.

10. The method according to any one of claims 7 to 9, characterized in that For each sample triple, the loss value corresponding to the sample triple is calculated according to the original encoding vector corresponding to the premise text sample, the positive encoding vector corresponding to the inherited text sample, and the negative encoding vector corresponding to the opposing text sample, including: For each of the sample triples, determining a first cosine similarity between an original encoding vector corresponding to the premise text sample and a positive example encoding vector corresponding to the inherited text sample; For each of the sample triples, determining M second cosine similarities between the original encoding vector corresponding to the premise text sample and the positive example encoding vectors corresponding to the M inherited text samples; For each of the sample triples, determining M third cosine similarities between the original encoding vector corresponding to the premise text sample and the negative example encoding vectors corresponding to the M opposing text samples; For each sample triplet, a loss value corresponding to the sample triplet is calculated according to the first cosine similarity, the M second cosine similarities, and the M third cosine similarities.

11. The method according to claim 1, characterized in that The method further comprises: Acquire a text sample set, wherein the text sample set includes N text samples and M sample triplets, each sample triplet includes a premise text sample, a successor text sample, and an opposing text sample, N is an integer greater than 1, and M is an integer greater than or equal to 1; Performing data enhancement processing on each text sample in the text sample set to obtain a positive example encoding vector corresponding to each text sample; Obtaining N negative example encoding vectors corresponding to each text sample in the text sample set, wherein the N negative example encoding vectors are generated according to the N text samples; According to the positive example encoding vector and the N negative example encoding vectors corresponding to each text sample, the model parameters of the text encoding model to be trained are updated to obtain a target text encoding model; For each sample triple, generate an original encoding vector corresponding to the premise text sample, a positive example encoding vector corresponding to the inherited text sample, and a negative example encoding vector corresponding to the opposing text sample; According to the original encoding vector corresponding to the premise text sample, the positive encoding vector corresponding to the inherited text sample and the negative encoding vector corresponding to the opposing text sample in each sample triplet, the target text encoding model is fine-tuned to obtain the text encoding model.

12. The method according to claim 1, characterized in that The step of generating a detection result for the account to be detected according to the K text vectors to be detected and the T abnormal text vectors includes: For each abnormal text vector in the T abnormal text vectors, calculate the similarity between the abnormal text vector and each text vector to be tested, and obtain K similarities, wherein the greater the similarity, the higher the similarity between the abnormal text vector and the text vector to be tested; For each abnormal text vector in the T abnormal text vectors, determine P test text vectors having the greatest similarity to the abnormal text vector according to the K similarities, wherein P is an integer greater than or equal to 1 and less than K; Determine the number of the K test texts containing abnormal content according to the P test text vectors corresponding to each abnormal text vector; A detection result for the account to be detected is generated according to the number of the texts to be detected that contain abnormal content.

13. The method according to claim 1, characterized in that The step of generating a detection result for the account to be detected according to the K text vectors to be detected and the T abnormal text vectors includes: Calculate the similarity between each of the K text vectors to be tested and each of the T abnormal text vectors to obtain K×T similarities, wherein the greater the similarity, the higher the similarity between the abnormal text vector and the text vector to be tested; If there is at least one similarity greater than or equal to the similarity threshold among the K×T similarities, then determining the number of the K test texts containing abnormal content according to the at least one similarity; A detection result for the account to be detected is generated according to the number of the texts to be detected that contain abnormal content.

14. The method according to claim 1, characterized in that The step of generating a detection result for the account to be detected according to the K text vectors to be detected and the T abnormal text vectors includes: Initialize Q cluster centers, where Q is an integer greater than 1; Calculate the Euclidean distance between each of the K test text vectors and each of the T abnormal text vectors and each cluster center to obtain (K+T)×Q Euclidean distances; According to the (K+T)×Q Euclidean distances, each of the text vectors to be tested and each of the abnormal text vectors are respectively divided into clusters corresponding to the cluster center with the smallest Euclidean distance, to obtain Q clusters; According to the Q clusters, each of the text vectors to be tested and each of the abnormal text vectors are divided until clustering optimization conditions are met, thereby obtaining Q target clusters; A detection result for the account to be detected is generated according to the number of abnormal text vectors included in each of the Q target clusters.

15. The method according to claim 1, characterized in that The step of generating a detection result for the account to be detected according to the K text vectors to be detected and the T abnormal text vectors includes: The K to-be-tested text vectors and the T abnormal text vectors are used as a text vector set, wherein the text vector set includes (K+T) text vectors; Selecting a text vector from the text vector set as a first text vector; Calculate the similarity between the first text vector and the remaining (K+T-1) text vectors in the text vector set, and obtain a second text vector corresponding to the maximum similarity; If the maximum similarity is greater than or equal to a target threshold, adding the second text vector to the first cluster corresponding to the first text vector, and updating the cluster center of the first cluster; If the maximum similarity is less than the target threshold, the second text vector is used as the cluster center of the second cluster; When each text vector in the text vector set is divided into a corresponding cluster, R clusters are obtained, wherein R is an integer greater than or equal to 1; A detection result for the account to be detected is generated according to the number of abnormal text vectors included in each of the R clusters.

16. The method according to claim 1, characterized in that After generating the detection result for the account to be detected according to the K text vectors to be detected and the T abnormal text vectors, the method further includes: When the detection result indicates that the account to be detected is a suspicious account, intercepting the text sent by the suspicious account; When the detection result indicates that the account to be detected is a malicious account, the text sent by the malicious account is intercepted and the usage permission of the malicious account is blocked.

17. A service processing device, characterized in that: include: An acquisition module, used for acquiring a text set to be tested, wherein the text set to be tested includes K texts to be tested sent through the account to be tested, and K is an integer greater than or equal to 1; An encoding module, used for encoding each test text in the test text set by using a text encoding model to obtain K test text vectors, wherein the K test text vectors have a one-to-one correspondence with the K test texts, and the text encoding model is obtained by comparative learning of positive example encoding vector pairs and negative example encoding vector pairs; The acquisition module is further used to acquire an abnormal text set, wherein the abnormal text set includes T abnormal texts marked as abnormal content, and T is an integer greater than or equal to 1; The encoding module is further used to encode each abnormal text in the abnormal text set using the text encoding model to obtain T abnormal text vectors, wherein the T abnormal text vectors have a one-to-one correspondence with the T abnormal texts; A generation module is used to generate a detection result for the account to be detected based on the K text vectors to be detected and the T abnormal text vectors, wherein the detection result is used to indicate the abnormality degree of the account to be detected.

18. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 16 are implemented.

19. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 16 are implemented.

20. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 16 are implemented.