Account identification method and device, computer equipment, storage medium and program product
By obtaining the current text and historical text of the account to be identified, and determining the similarity model based on the text variant direction of the historical violation text, calculating the text similarity to identify the violation account, the problem of low accuracy of account identification in the existing technology is solved, and efficient and accurate identification of screen-swimming violations is achieved.
Patent Information
- Application Number
- CN202510093508.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-06
AI Technical Summary
The existing technology has low accuracy when identifying accounts for screen-swimming violations, making it difficult to effectively identify complex violations.
By obtaining the current text of the account to be identified and its adjacent continuous historical text, a dynamic context analysis basis is established; at the same time, based on the text variant directions of multiple historical violation texts, the target similarity model is determined from multiple similarity models, and the similarity between the historical text and the current text is calculated through the model to determine whether the account to be identified belongs to the preset type account.
It improves the ability to capture complex violations, realizes efficient and accurate identification of screen-swiping violations, and improves the accuracy of account identification.
Smart Images

Figure CN120104798A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of natural language processing, and in particular to an account identification method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] As the user scale of social platforms continues to expand, a large number of illegal behaviors such as spamming have gradually appeared on social platforms. For example, the same user account publishes various suspected advertisements, traffic diversion, fraud, incitement and other content on various channels of the platform. The prevalence of these illegal behaviors not only occupies a large amount of platform resources and affects the normal user experience, but may also cause excessive server load or even response delays. Therefore, it is necessary to identify these illegal contents and illegal accounts.
[0003] Traditional technologies usually rely on keyword matching or full-text matching rules to detect illegal content and identify illegal accounts. However, current account identification methods can only target accounts that post simple, repetitive illegal content, and the accuracy of account identification is low. Summary of the invention
[0004] Based on this, it is necessary to provide an account identification method, apparatus, computer device, computer-readable storage medium and computer program product that can improve the accuracy of account identification in response to the above technical problems.
[0005] In a first aspect, the present application provides an account identification method, comprising:
[0006] Acquire a current text input by the account to be identified through the social platform and a preset number of historical texts; the historical texts are texts continuously input by the account to be identified before the current text;
[0007] Determining a target similarity model from a plurality of similarity models according to text variation directions of a plurality of historical illegal texts; the historical illegal texts comprising texts inputted through the social platform by at least one account of a preset type within a preset historical time period;
[0008] Determine the similarity between each of the historical texts and the current text by using the target similarity model;
[0009] According to the similarity, it is determined whether the account to be identified belongs to the preset type of account.
[0010] In one embodiment, determining a target similarity model from a plurality of similarity models according to the text variation directions of a plurality of historical illegal texts includes:
[0011] Determining the text variation direction with the highest proportion from the text variation directions of the plurality of historical illegal texts;
[0012] The target similarity model is determined from a plurality of similarity models according to the text variation direction with the highest proportion.
[0013] In one embodiment, the target similarity model includes a literal similarity model; and determining the similarity between each of the historical texts and the current text by using the target similarity model includes:
[0014] Obtain the string lengths of the historical text and the current text;
[0015] When the length of the character string does not satisfy the calculation termination condition, the edit distance between each of the historical texts and the current text is calculated by the character similarity model, and the similarity between each of the historical texts and the current text is determined according to the edit distance;
[0016] The edit distance is used to represent the minimum number of editing operations required to convert the historical text into the current text.
[0017] In one of the embodiments, after obtaining the character string lengths of the historical text and the current text, the method further includes:
[0018] When the length of the character string meets the calculation termination condition, a preset value is returned;
[0019] The similarity between each of the historical texts and the current text is determined according to the preset value through the word similarity model.
[0020] In one embodiment, the target similarity model includes a semantic similarity model; and determining the similarity between each of the historical texts and the current text by using the target similarity model includes:
[0021] Obtaining a first text vector corresponding to each of the historical texts and a second text vector corresponding to the current text;
[0022] The cosine distance between each of the first text vectors and the second text vector is calculated by the semantic similarity model, and the similarity between each of the historical texts and the current text is determined according to the cosine distance.
[0023] In one embodiment, determining whether the account to be identified belongs to the preset type of account according to the similarity includes:
[0024] Determine the target text whose similarity is greater than the similarity threshold from each of the historical texts;
[0025] According to the text quantity of the target text and the statistical information of each similarity, it is determined whether the account to be identified belongs to the preset type of account.
[0026] In a second aspect, the present application further provides an account identification device, including:
[0027] An acquisition module, used to acquire a current text input by the account to be identified through the social platform and a preset number of historical texts; the historical texts are texts continuously input by the account to be identified before the current text;
[0028] A model determination module, configured to determine a target similarity model from a plurality of similarity models according to text variation directions of a plurality of historical illegal texts; the historical illegal texts comprising texts inputted through the social platform by at least one account of a preset type within a preset historical time period;
[0029] A similarity determination module, used to determine the similarity between each of the historical texts and the current text through the target similarity model;
[0030] The account determination module is used to determine whether the account to be identified belongs to the preset type of account according to the similarity.
[0031] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0032] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above method when executed by a processor.
[0033] In a fifth aspect, the present application also provides a computer program product, including a computer program, which implements the steps of the above method when executed by a processor.
[0034] The above-mentioned account identification method, device, computer equipment, computer-readable storage medium and computer program product obtain the current text input by the account to be identified through the social platform and a preset number of historical texts, wherein the historical text is the text continuously input by the account to be identified before the current text; determine the target similarity model from multiple similarity models according to the text variation directions of multiple historical violation texts, wherein the historical violation text includes text input by at least one preset type account through the social platform within a preset historical time period; determine the similarity between each historical text and the current text through the target similarity model; and determine whether the account to be identified belongs to the preset type account according to the similarity. By obtaining the current text of the account to be identified and its adjacent continuous historical text, the foundation for contextual dynamic analysis is established; at the same time, the target similarity model is determined by the text variation directions of multiple historical illegal texts, which can adapt to the changing characteristics of different types of illegal texts and improve the ability to capture complex illegal content; by using the target similarity model to calculate the similarity between historical text and current text, screen-swiping illegal behaviors can be identified efficiently and accurately; by judging whether the account to be identified belongs to the preset type of account based on the similarity, accurate identification of user account behavior can be achieved, thereby improving the accuracy of account identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0036] Figure 1 A diagram of an application environment of an account identification method in an embodiment;
[0037] Figure 2 A flowchart of an account identification method in an embodiment;
[0038] Figure 3 A flowchart of an account identification method in another embodiment;
[0039] Figure 4 It is a structural block diagram of an account identification device in one embodiment;
[0040] Figure 5 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0042] The account identification method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal 102 obtains the current text entered by the account to be identified through the social platform and a preset number of historical texts; the historical text is the text continuously entered by the account to be identified before the current text; the terminal 102 determines the target similarity model from multiple similarity models according to the text variation direction of multiple historical violation texts; the historical violation text includes at least one text entered by a preset type account through the social platform within a preset historical time period; the terminal 102 determines the similarity between each historical text and the current text through the target similarity model; the terminal 102 determines whether the account to be identified belongs to the preset type account according to the similarity. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices, and the Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car devices, projection devices, etc. The portable wearable device may be a smart watch, a smart bracelet, a head-mounted device, etc. The head-mounted device may be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server 104 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0043] In an exemplary embodiment, Figure 2 As shown, an account identification method is provided, which is applied to Figure 1 The terminal 102 in FIG. 1 is used as an example to illustrate, including:
[0044] Step S202, obtaining the current text and a preset number of historical texts input by the account to be identified through the social platform.
[0045] In a specific implementation, the account to be identified may be any user account registered on the social platform; or, the account to be identified may also be any user account with a high level of activity on the social platform, such as a user account with a high recent login frequency and a high frequency of posting social content.
[0046] Among them, the social platform can be a social platform mainly in the form of multimedia such as text, pictures, videos, and voice, which can realize instant communication and high-frequency interaction. Optionally, the social platform can have real-time voice function, video live broadcast function, and can support online chat rooms or voice rooms, as well as support sending dynamic messages or notifications.
[0047] The current text is the text content input or published by the account to be identified at the current time point, in other words, the latest text content published by the account to be identified.
[0048] The historical texts are the texts that were continuously input by the account to be identified before the current text. These historical texts are input in chronological order without skipping the intermediate inputs. In other words, these historical texts are inputs that are close in time, rather than jumping inputs.
[0049] Among them, the number of historical texts meets the preset number. In a specific implementation, the preset number can be a value determined based on actual business needs, for example, it can be 30 historical texts; or, the preset number can also be a value dynamically adjusted according to the activity or behavior pattern of the user account. Optionally, the preset number of historical texts can use a sliding window strategy, for example, the size of the sliding window is set to k, k is a preset number, and the most recent k historical texts can be obtained through the sliding window. The more preset numbers of historical texts, the higher the accuracy of text analysis, but the computational complexity will also increase.
[0050] In one embodiment, obtaining a current text input by the account to be identified through a social platform and a preset number of historical texts includes: obtaining a plurality of original texts input by the account to be identified through the social platform; standardizing the character format and text length of the plurality of original texts to obtain a plurality of standardized texts; and determining the current text input by the account to be identified through the social platform and a preset number of historical texts from the plurality of standardized texts.
[0051] In the specific implementation, the character formats of multiple original texts are standardized, which may include removing spaces and invisible characters, lowercase letters, converting traditional and simplified Chinese characters, converting to half-width characters, standardizing symbols, and other operations; the text lengths of multiple original texts are standardized, which may include operations such as text length truncation, and other operations. Among them, removing spaces and invisible characters can be the identification and removal of various spaces and invisible characters; lowercase letters can be the conversion of all letters to lowercase; traditional-simplified conversion can be the identification of traditional Chinese characters and their conversion to simplified Chinese; half-width conversion can be the uniform half-width conversion of non-half-width characters in the text; symbol standardization can be the identification and standardization of communication identifiers such as mobile phone numbers, while maintaining the original length, replacing each number with a unified symbol, such as the "#" sign; text length truncation can be the standardized truncation of overly long text, for example, only the first 30 characters are retained after truncation. This is because the illegal user account usually uses highly repetitive text content when swiping the screen, and the core content is usually concentrated in the front part of the text. The extra length of the text is not very important, so the text length can be truncated to improve the efficiency of subsequent similarity calculations; and the text length truncation operation can also effectively defend against maliciously constructed over-long strings that attempt to exhaust system resources.
[0052] The technical solution of the above-mentioned embodiment can standardize the text format as much as possible, improve the efficiency of subsequent similarity calculation, and thus improve the efficiency and accuracy of identifying screen-swiping violations.
[0053] In actual applications, the account to be identified can input text through various channels of the social platform, such as posts, comments, private messages, voice room chats, etc. If the account to be identified sends illegal text such as screen swiping, the text input in the same channel is generally highly repetitive. Therefore, the terminal obtains the current text and a preset number of historical texts input by the account to be identified through the social platform, which can be the current text and a preset number of historical texts input by the account to be identified through any channel of the social platform. Alternatively, the terminal can detect the most active channel of the account to be identified on the social platform, and obtain the current text and a preset number of historical texts input by the account to be identified through the most active channel. The most active channel can be the channel with the highest number of texts input by the account to be identified in the recent time period, or the channel with the highest frequency of use by the account to be identified in the recent time period.
[0054] Step S204, determining a target similarity model from a plurality of similarity models according to the text variation directions of a plurality of historical illegal texts.
[0055] The historical violation text includes text entered by at least one preset type of account through the social platform within a preset historical time period.
[0056] Among them, historical illegal texts can be texts that have been identified as meeting the illegal behavior on social platforms, such as texts used to carry out network attacks, steal information, extortion fraud, steal money and other network illegal behaviors. These texts can also be called black industry texts, that is, texts related to the network black industry chain. Texts that meet the illegal behavior can be screen-sweeping information with a high repetition rate, including advertisements, diversion, fraud, incitement and other text content.
[0057] Among them, the preset type account can be a violation account, and the behavior pattern of the preset type account is closely related to the violation. The preset type account usually enters text that conforms to the violation, or enters black market text. The preset historical time period can be a recent time period, such as the past month, the past 7 days, etc.
[0058] The text variation direction refers to the changing trend and characteristics of text content in different scenarios or time periods, which is used to describe the change pattern of text at the formal and semantic levels. Its changes are usually driven by specific intentions, such as evading detection, increasing diversity, or hiding illegal features.
[0059] In a specific implementation, the terminal may identify the text variant directions of multiple historical illegal texts in the following ways: for the literal variant direction, the difference in text form may be captured by detecting explicit feature changes such as character replacement, insertion of interference characters, and vocabulary adjustment; for the semantic variant direction, a semantic embedding model may be used to calculate the semantic similarity between texts, and combined with syntactic structure analysis, the situation where the expression method changes but the meaning is the same may be identified, thereby determining the text variant direction of each historical illegal text.
[0060] In a specific implementation, the terminal may obtain the historical violation texts input by each recently identified violation account, based on the text variation direction of multiple historical violation texts, or the terminal may obtain the historical violation texts input by the recently identified serious violation account, based on the text variation direction of multiple historical violation texts. For example, the terminal may classify the severity of the violation, such as advertising screen swiping is a medium violation, inciting screen swiping or fraudulent screen swiping is a serious violation, and then obtain the historical violation texts input by the serious violation account.
[0061] In a specific implementation, the terminal determines a target similarity model from multiple similarity models according to the text variation directions of multiple historical illegal texts, and may determine the target similarity model from multiple similarity models according to the proportion of the text variation directions of multiple historical illegal texts.
[0062] Exemplarily, the terminal can determine the text variant direction with the highest proportion based on the proportions of the text variant directions of multiple historical violation texts, and then determine the target similarity model that matches the text variant direction with the highest proportion from multiple similarity models; or, when the proportion difference between the proportions of each text variant direction is less than the difference threshold, the similarity models that match each text variant direction can be jointly determined as the target similarity model from multiple similarity models.
[0063] In one embodiment, the text variant direction may include a literal variant direction and a semantic variant direction.
[0064] Among them, the literal variant direction may refer to the explicit change of the text at the character level or the lexical level without changing or only partially changing the semantics of the text. For example, replacing with similar characters, such as replacing "free" with "免¥费" or "mian费"; inserting irrelevant characters, such as replacing "time-limited discount" with "限#时优$惠"; adjusting the order of words or phrases, such as replacing "low-price promotion" with "promotion low price", etc.
[0065] Among them, the semantic variant direction may refer to the implicit change that occurs at the semantic level of the text, that is, by modifying the expression way to change the literal expression while retaining the core meaning or intention. For example, replacing with synonyms with similar semantics, such as replacing "purchase" with "place an order", "free trial" with "experience without payment"; changing the grammatical structure or sentence pattern, such as replacing "Our store has a time-limited discount" with "The time-limited discount is available in our store"; using non-specific language or vague expressions, such as replacing "Come and snap it up" with "The opportunity is rare, act quickly".
[0066] Among them, the similarity model is an algorithm model used to evaluate the similarity between two texts, and the similarity model may include a literal similarity model, a semantic similarity model, etc. For example, the literal similarity model may be a literal similarity model based on the edit distance, a literal similarity model based on the Jaccard similarity, etc., for determining the similarity of literal variant texts; the semantic similarity model may be a semantic embedding model based on deep learning, such as the BERT model, the Word2Vec model, etc., for determining the similarity of semantic variant texts.
[0067] The target similarity model is the model that is most suitable for determining the similarity between the historical text and the current text among multiple similarity models. Dynamically selecting the target similarity model according to the text variant direction of the historical violation text can ensure that the target similarity model adapts to the specific characteristics of the violation text.
[0068] Step S206, determine the similarity between each historical text and the current text through the target similarity model.
[0069] Among them, similarity can be used to indicate the similarity between two texts in terms of content, semantics or form, and can be expressed by a quantitative index (such as a value between 0 and 1), for example, 1 indicates that they are exactly the same, and 0 indicates that they are completely different. Optionally, the target similarity model can calculate the similarity through the following indicators: edit distance, Jaccard similarity coefficient, cosine similarity, etc.
[0070] In a specific implementation, the terminal determines the similarity between each historical text and the current text through a target similarity model, and the similarity between each historical text and the current text is calculated one by one.
[0071] For example, the input of the target similarity model can be expressed as: {text: y, previous_text_list:[x 1 , x 2 , …, x n ], threshold: a}. Among them, "text" can be represented as the current text, with content y; "previous_text_list" can be represented as a preset number of historical texts, respectively x 1 、x 2 ,……,x n etc.; "threshold" can represent the similarity threshold, which can be used to compare with the similarity as a benchmark for judging the size of the similarity. For example, the similarity threshold can be 0.7 or 0.8. In the specific implementation, you can first determine x 1 Similarity with y, then determine x 2 The similarity with y, and so on, until x is determined n Similarity with y.
[0072] Step S208: Determine whether the account to be identified belongs to a preset type of account based on the similarity.
[0073] In a specific implementation, the terminal may compare the similarity with a similarity threshold. Assuming that the similarity between any historical text and the current text is greater than the similarity threshold, it may be determined that the historical text is a similar text. Assuming that the number of similar texts in each historical text is greater than the quantity threshold, it may be determined that the account to be identified has entered a large number of similar texts with a high degree of repetition, thereby determining that the account to be identified has screen-swiping violations, and determining the account to be identified as a preset type account, that is, a violation account; alternatively, the terminal may also calculate the average value of the similarities between each historical text and the current text to obtain an average similarity. If the average similarity is greater than a certain threshold, it may be determined that the account to be identified is a preset type account.
[0074] In actual applications, after determining the preset type of account, the terminal can block the screen-swiping violation behavior of the preset type of account in real time, for example, it can block the text input by the preset type of account, prohibit the preset type of account from inputting text, block the preset type of account, etc. Optionally, the terminal can determine the severity level of the preset type of account according to the violation type of the preset type of account, and determine the corresponding blocking measures according to the severity level. For example, if the fraud type is a high severity level, the preset type of account can be directly blocked; if the advertising type is a low severity level, the text input by the preset type of account can be blocked.
[0075] In practical applications, text recall can refer to the process of screening out texts related to the target text from a text collection based on specific standards or algorithms. In this application, text recall can refer to finding texts with high similarity to the current text from historical texts as much as possible by calculating text similarity, thereby realizing text recall of illegal texts and further realizing the identification of preset types of accounts with illegal behaviors.
[0076] In the above account identification method, the current text input by the account to be identified through the social platform and a preset number of historical texts are obtained, wherein the historical text is the text continuously input by the account to be identified before the current text; according to the text variation direction of multiple historical violation texts, a target similarity model is determined from multiple similarity models, wherein the historical violation text includes text input by at least one preset type account through the social platform within a preset historical time period; the similarity between each historical text and the current text is determined by the target similarity model; according to the similarity, it is determined whether the account to be identified belongs to the preset type account. By obtaining the current text of the account to be identified and its adjacent continuous historical text, the basis for contextual dynamic analysis is established; at the same time, by determining the target similarity model through the text variation direction of multiple historical violation texts, it can adapt to the changing characteristics of different types of violation texts and improve the ability to capture complex violation content; by using the target similarity model to calculate the similarity between the historical text and the current text, the screen-swiping type violation behavior can be efficiently and accurately identified; based on the similarity, whether the account to be identified belongs to the preset type account can be accurately identified for the user account behavior, thereby improving the accuracy of account identification.
[0077] In another embodiment, a target similarity model is determined from multiple similarity models based on text variation directions of multiple historical illegal texts, including: determining the text variation direction with the highest proportion from the text variation directions of multiple historical illegal texts; and determining the target similarity model from multiple similarity models based on the text variation direction with the highest proportion.
[0078] For example, assuming that among the text variant directions of multiple historical illegal texts, the literal variant direction accounts for 70% and the semantic variant direction accounts for 30%, the text variant direction with the highest proportion is the literal variant direction, and the literal variant similarity model can be determined as the target similarity model.
[0079] In one embodiment, the text variation direction includes a literal variation direction and a semantic variation direction; determining a target similarity model from multiple similarity models according to the text variation direction with the highest proportion may include:
[0080] When the proportion corresponding to the text variant direction with the highest proportion is greater than the proportion threshold, a target similarity model matching the text variant direction with the highest proportion can be determined from multiple similarity models; when the proportion corresponding to the text variant direction with the highest proportion is less than or equal to the proportion threshold, a literal similarity model and a semantic similarity model can be determined as the target similarity model from multiple similarity models.
[0081] Then, the model weights of the literal similarity model and the semantic similarity model can be determined based on the proportion of the literal variant direction and the semantic variant direction; the similarity between each historical text and the current text can be determined through the target similarity model, which can be that the first similarity between each historical text and the current text is calculated through the literal similarity model, and the second similarity between each historical text and the current text is calculated through the semantic similarity model; the similarity between each historical text and the current text is obtained according to the model weight of the first similarity and the literal similarity model, and the model weight of the second similarity and the semantic similarity model.
[0082] The technical solution of this embodiment can achieve accurate matching for different types of text variants by dynamically identifying the main variant directions of recent historical illegal texts and flexibly selecting the most suitable similarity model. It can also flexibly respond to the variant trends and evolution directions of illegal texts, and timely adjust the text similarity detection strategy to improve the accuracy and flexibility of text similarity detection, thereby improving the accuracy of identifying preset types of accounts.
[0083] In another embodiment, the target similarity model includes a literal similarity model; determining the similarity between each historical text and the current text through the target similarity model includes: obtaining the string length of the historical text and the current text; when the string length does not meet the calculation termination condition, calculating the edit distance between each historical text and the current text through the literal similarity model, and determining the similarity between each historical text and the current text based on the edit distance.
[0084] The word similarity model may be a model for calculating the degree of similarity between texts in words, for example, a word similarity model based on edit distance, Jaccard similarity, or N-gram matching.
[0085] The calculation termination condition is used to determine whether the edit distance needs to be calculated. In the specific implementation, if the string length does not meet the calculation termination condition, the edit distance needs to be calculated; if the string length meets the calculation termination condition, the calculation of the edit distance needs to be terminated.
[0086] In one embodiment, the calculation termination condition includes: the difference in character string length between the historical text and the current text exceeds a first threshold, or the character string length of the historical text or the current text is less than a second threshold.
[0087] The first threshold is used to limit the maximum length difference between two character strings, because when the length difference between two character strings is large, the similarity between them may be very low, and there is no need to perform a complete edit distance calculation. For example, the first threshold can be 10 or 20, which is not specifically limited. The second threshold can be a value close to 0. Assuming that the string length of the historical text or the current text is 0, since the similarity between a string with a length of 0 and any other string should be the lowest, there is no need to further calculate the edit distance. Therefore, if the string length meets the calculation termination condition, it can be further explained that the similarity between the historical text and the current text is very low.
[0088] The edit distance is used to characterize the minimum number of edit operations required to convert a historical text into the current text. For example, the edit distance can be the Levenshtein distance, which refers to the minimum number of operations required to convert one string into another string through the three operations of insertion, deletion, and substitution. For example, the edit distance between hello and h3llo is 1. The smaller the edit distance between two texts, the higher the similarity between the two texts.
[0089] In a specific implementation, the terminal may determine the similarity between each historical text and the current text according to the edit distance, wherein the smaller the edit distance is, the higher the similarity between the historical text and the current text is.
[0090] For example, the similarity can be expressed as:
[0091] similarity=(10-score) / 1; (1)
[0092] Among them, score is the edit distance, similarity is the similarity, and the edit distance can be converted to similarity through the linear relationship in the above formula. It can be seen that the smaller the edit distance, the closer the similarity is to 1, and the higher the similarity; the larger the edit distance, the closer the similarity is to 0, and the lower the similarity.
[0093] The technical solution of this embodiment reduces meaningless edit distance calculations, saves computing resources, and improves overall processing efficiency by judging whether the length of a character string meets the calculation termination condition. The calculation of the edit distance is triggered only when the calculation termination condition is not met, thereby accurately and efficiently measuring the similarity between texts by calculating the edit distance.
[0094] In another embodiment, after obtaining the string lengths of the historical texts and the current text, it also includes: returning a preset value when the string length meets the calculation termination condition; and determining the similarity between each historical text and the current text according to the preset value through a literal similarity model.
[0095] The preset value may be a fixed value directly returned when the string length meets the termination condition. If the string length meets the calculation termination condition, it can further indicate that the similarity between the historical text and the current text is very low, so the similarity determined based on the returned preset value is also very low.
[0096] As an example, there may be a certain correspondence between the preset value and the similarity. For example, if the preset value is 10, when the preset value is 10, the similarity may be equal to 0, which is the lowest similarity.
[0097] As an example, if the similarity is calculated based on formula (1), for example, the preset value is 10, the preset value can be regarded as the edit distance score and substituted into formula (1), and it can be seen that the similarity is equal to 0, which is the lowest similarity.
[0098] The technical solution of this embodiment directly returns a preset value when the string length meets the calculation termination condition, without further complex similarity calculation, thereby reducing meaningless calculations and significantly improving the efficiency of text detection and account recognition.
[0099] In another embodiment, the target similarity model includes a semantic similarity model; determining the similarity between each historical text and the current text through the target similarity model includes: obtaining a first text vector corresponding to each historical text and a second text vector corresponding to the current text; calculating the cosine distance between each first text vector and the second text vector through the semantic similarity model, and determining the similarity between each historical text and the current text based on the cosine distance.
[0100] Among them, the semantic similarity model can be a model used to measure the similarity of texts at the semantic level, which can understand the semantic meaning of the text rather than just the literal content.
[0101] The text vector refers to a high-dimensional vector generated after the text is mapped to the vector space through a semantic embedding model. Each dimension of the vector reflects the specific semantic representation of the text. The first text vector is the text vector corresponding to the historical text, and the second text vector is the text vector of the current text.
[0102] Among them, the cosine distance can be used to measure the similarity of the angle between two text vectors in space. The value of the cosine distance is usually between [0, 1]. The smaller the value of the cosine distance, the higher the similarity. The cosine distance can be expressed as:
[0103] ;
[0104] Here, u and v may represent the first text vector and the second text vector respectively.
[0105] The technical solution of this embodiment achieves efficient and accurate judgment of the semantic similarity between historical text and current text through the calculation of semantic similarity model and cosine distance, thereby improving the accuracy of account recognition.
[0106] In another embodiment, determining whether the account to be identified belongs to a preset type of account based on similarity includes: determining a target text whose similarity is greater than a similarity threshold from each historical text; and determining whether the account to be identified belongs to a preset type of account based on the text quantity of the target text and statistical information of each similarity.
[0107] The similarity threshold can be used to determine whether two texts are similar enough. Only when the similarity is greater than this threshold, the two texts are considered similar. For example, if the similarity is in the range of [0,1], the similarity threshold can be 0.7 or 0.8.
[0108] In the specific implementation, the similarity threshold can be a value set by the user according to business needs, which provides flexibility for business logic. The terminal can adjust the similarity threshold according to different application scenarios. Only when the similarity between the historical text and the current text is greater than the similarity threshold, can the historical text and the current text be considered similar. By flexibly adjusting the accurate similarity threshold, effective identification of text variant behaviors can be achieved, thereby achieving efficient and accurate text recognition and account identification.
[0109] For example, when the current application scenario is a high-risk scenario, for example, there are recent serious violations such as fraud with a high frequency, the similarity threshold can be lowered to increase the sensitivity to screen-sweeping illegal texts; when the current application scenario is a low-risk scenario, for example, there are only minor violations such as advertising screen-sweeping in the recent period and the frequency of the behaviors is low, the similarity threshold can be increased to reduce the misjudgment rate.
[0110] The target text may be a text selected from historical texts, whose similarity with the current text exceeds a similarity threshold. The number of target texts may be used to reflect the number of texts in historical texts that are similar to the current text. Optionally, if the number of target texts exceeds a number threshold, it may be determined that there are enough texts in multiple historical texts that are similar to the current text, indicating that the historical texts and the current texts input by the account to be identified are consistent with the characteristics of screen-swiping illegal texts, thereby further indicating that the account to be identified belongs to a preset type of account.
[0111] The statistical information of each similarity may include the average score, the highest score, and the lowest score of each similarity. If the average score, the highest score, and the lowest score of each similarity are all very high, it can be explained that the account to be identified has entered a large number of similar texts, which is consistent with the characteristics of the illegal text of the screen spamming type, and thus it can be further explained that the account to be identified belongs to the preset type account.
[0112] In a specific implementation, the terminal determines whether the account to be identified belongs to a preset type of account based on the amount of text in the target text and the statistical information of each similarity. The terminal may determine that the account to be identified belongs to a preset type of account when the amount of text in the target text is greater than the amount threshold and the average score of each similarity is higher than the average score threshold; or, the terminal may determine that the account to be identified belongs to a preset type of account when the amount of text in the target text is greater than the amount threshold, the average score of each similarity is higher than the average score threshold, and the highest score of each similarity is greater than the highest score threshold.
[0113] Since the terminal obtains the current text input by the account to be identified through the social platform and a preset number of historical texts, it is possible to first obtain multiple original texts input by the account to be identified through the social platform; standardize the character format and text length of the multiple original texts to obtain multiple standardized texts; and then determine the current text input by the account to be identified through the social platform and a preset number of historical texts from the multiple standardized texts.
[0114] Exemplarily, the output of the target similarity model may include: {Similarity calculation details list: [{Original text: xxx, Standardized text: xxx, Similarity: xxx, Is it similar: true / false}, ...]}. The Standardized text includes the standardized current text and a preset number of historical texts. The similarity is the similarity between each historical text and the current text. If the similarity is greater than the similarity threshold, it can be determined that the historical text is similar to the current text, and the result is true; if the similarity is less than the similarity threshold, it can be determined that the historical text is not similar to the current text, and the result is false.
[0115] Exemplarily, the output of the target similarity model may also include: {average score: xxx, highest score: xxx, lowest score: xxx, number of target texts: xxx}, which is used to represent the number of target texts and statistical information of each similarity.
[0116] The technical solution of this embodiment, through similarity calculation and threshold filtering, accurately screens out target texts that are highly relevant to the account to be identified, avoids the interference of low-relevant texts on the judgment results, and improves the accuracy of classifying the account to be identified; through multi-dimensional data such as the number of target texts and similarity statistical information, it comprehensively extracts account behavior characteristics, can effectively capture the behavioral patterns of illegal accounts, improve the coverage of illegal behavior detection and the accuracy of preset type account identification.
[0117] In another embodiment, Figure 3 As shown, an account identification method is provided, which is applied to Figure 1 The terminal 102 in the example is used as an example to illustrate, and the following steps are included:
[0118] Step S302, obtaining the current text and a preset number of historical texts input by the account to be identified through the social platform.
[0119] The historical text is the text that was continuously input by the account to be identified before the current text.
[0120] Step S304, determining the text variation direction with the highest proportion from the text variation directions of the plurality of historical illegal texts.
[0121] The historical violation text includes text input by at least one account of a preset type through the social platform within a preset historical time period.
[0122] Step S306, determining a target similarity model from multiple similarity models according to the text variation direction with the highest proportion; if the target similarity model is a literal similarity model, executing step S308; if the target similarity model is a semantic similarity model, executing step S316.
[0123] Step S308, obtaining the character string lengths of the historical text and the current text.
[0124] Step S310, determine whether the character string length meets the calculation termination condition, if not, execute step S312; if yes, execute step S314.
[0125] The calculation termination conditions include: the difference in character string length between the historical text and the current text exceeds a first threshold, or the character string length of the historical text or the current text is less than a second threshold.
[0126] Step S312, calculating the edit distance between each historical text and the current text through the word similarity model, and determining the similarity between each historical text and the current text according to the edit distance, and executing step S320.
[0127] Among them, the edit distance is used to represent the minimum number of editing operations required to convert the historical text into the current text.
[0128] Step S314, returning a preset value, and determining the similarity between each historical text and the current text according to the preset value through the word similarity model, and executing step S320.
[0129] Step S316, obtaining the first text vector corresponding to each historical text and the second text vector corresponding to the current text.
[0130] Step S318, calculating the cosine distance between each first text vector and the second text vector through the semantic similarity model, and determining the similarity between each historical text and the current text according to the cosine distance, and executing step S320.
[0131] Step S320: determining a target text whose similarity is greater than a similarity threshold from each historical text.
[0132] Step S322, determining whether the account to be identified belongs to a preset type of account based on the text quantity of the target text and the statistical information of each similarity.
[0133] It should be noted that the specific limitations of the above steps can refer to the specific limitations of an account identification method above.
[0134] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0135] Based on the same inventive concept, the embodiment of the present application also provides an account identification device for implementing the account identification method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more account identification device embodiments provided below can refer to the limitations on the account identification method above, and will not be repeated here.
[0136] In an exemplary embodiment, Figure 4 As shown, an account identification device is provided, comprising:
[0137] The acquisition module 410 is used to acquire the current text input by the account to be identified through the social platform and a preset number of historical texts; the historical texts are the texts continuously input by the account to be identified before the current text.
[0138] The model determination module 420 is used to determine a target similarity model from multiple similarity models according to the text variation directions of multiple historical illegal texts; the historical illegal texts include texts input by at least one preset type of account through the social platform within a preset historical time period.
[0139] The similarity determination module 430 is used to determine the similarity between each of the historical texts and the current text through the target similarity model.
[0140] The account determination module 440 is used to determine whether the account to be identified belongs to the preset type of account according to the similarity.
[0141] In one embodiment, the model determination module 420 is specifically used to determine the text variant direction with the highest proportion from the text variant directions of the multiple historical illegal texts; and determine the target similarity model from multiple similarity models based on the text variant direction with the highest proportion.
[0142] In one embodiment, the target similarity model includes a literal similarity model; the similarity determination module 430 is specifically used to obtain the string lengths of the historical text and the current text; when the string length does not meet the calculation termination condition, the edit distance between each of the historical texts and the current text is calculated by the literal similarity model, and the similarity between each of the historical texts and the current text is determined based on the edit distance; wherein the edit distance is used to characterize the minimum number of editing operations for converting the historical text into the current text.
[0143] In one embodiment, the similarity determination module 430 is specifically used to return a preset value when the character string length meets the calculation termination condition; and determine the similarity between each of the historical texts and the current text according to the preset value through the literal similarity model.
[0144] In one embodiment, the calculation termination condition includes: the difference in character string length between the historical text and the current text exceeds a first threshold, or the character string length of the historical text or the current text is less than a second threshold.
[0145] In one embodiment, the target similarity model includes a semantic similarity model; the similarity determination module 430 is specifically used to obtain the first text vector corresponding to each of the historical texts and the second text vector corresponding to the current text; calculate the cosine distance between each of the first text vectors and the second text vector through the semantic similarity model, and determine the similarity between each of the historical texts and the current text based on the cosine distance.
[0146] In one embodiment, the account determination module 440 is specifically used to determine the target text whose similarity is greater than the similarity threshold from each of the historical texts; and determine whether the account to be identified belongs to the preset type of account based on the text quantity of the target text and the statistical information of each similarity.
[0147] Each module in the above account identification device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.
[0148] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 5As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, an account identification method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.
[0149] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0150] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0151] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0152] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0153] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0154] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0155] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0156] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be construed as limiting the scope of the present application. It should be noted that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. An account identification method, characterized in that: The method comprises: Acquire a current text input by the account to be identified through the social platform and a preset number of historical texts; the historical texts are texts continuously input by the account to be identified before the current text; Determining a target similarity model from a plurality of similarity models according to text variation directions of a plurality of historical illegal texts; the historical illegal texts comprising texts inputted through the social platform by at least one account of a preset type within a preset historical time period; Determine the similarity between each of the historical texts and the current text by using the target similarity model; According to the similarity, it is determined whether the account to be identified belongs to the preset type of account.
2. The method according to claim 1, characterized in that The method of determining a target similarity model from a plurality of similarity models according to the text variation directions of a plurality of historical illegal texts includes: Determining the text variation direction with the highest proportion from the text variation directions of the plurality of historical illegal texts; The target similarity model is determined from a plurality of similarity models according to the text variation direction with the highest proportion.
3. The method according to claim 1, characterized in that: The target similarity model includes a literal similarity model; and determining the similarity between each of the historical texts and the current text by using the target similarity model includes: Obtain the string lengths of the historical text and the current text; When the length of the character string does not satisfy the calculation termination condition, the edit distance between each of the historical texts and the current text is calculated by the character similarity model, and the similarity between each of the historical texts and the current text is determined according to the edit distance; The edit distance is used to represent the minimum number of editing operations required to convert the historical text into the current text.
4. The method according to claim 3, characterized in that After obtaining the character string lengths of the historical text and the current text, the method further includes: When the length of the character string meets the calculation termination condition, a preset value is returned; The similarity between each of the historical texts and the current text is determined according to the preset value through the word similarity model.
5. The method according to claim 1, characterized in that The target similarity model includes a semantic similarity model; and determining the similarity between each of the historical texts and the current text by using the target similarity model includes: Obtaining a first text vector corresponding to each of the historical texts and a second text vector corresponding to the current text; The cosine distance between each of the first text vectors and the second text vector is calculated by the semantic similarity model, and the similarity between each of the historical texts and the current text is determined according to the cosine distance.
6. The method according to claim 1, characterized in that The step of determining whether the account to be identified belongs to the preset type of account according to the similarity includes: Determine the target text whose similarity is greater than the similarity threshold from each of the historical texts; According to the text quantity of the target text and the statistical information of each similarity, it is determined whether the account to be identified belongs to the preset type of account.
7. An account identification device, characterized in that: The device comprises: An acquisition module, used to acquire a current text input by the account to be identified through the social platform and a preset number of historical texts; the historical texts are texts continuously input by the account to be identified before the current text; A model determination module, configured to determine a target similarity model from a plurality of similarity models according to text variation directions of a plurality of historical illegal texts; the historical illegal texts comprising texts inputted through the social platform by at least one account of a preset type within a preset historical time period; A similarity determination module, used to determine the similarity between each of the historical texts and the current text through the target similarity model; The account determination module is used to determine whether the account to be identified belongs to the preset type of account according to the similarity.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.