Black market information identification method, device and equipment, storage medium and program product

By combining a large language model and the isolated forest algorithm, and using a smoothed exponential-logarithmic asymmetric unit activation function to train the model, the problem of low efficiency in identifying black and gray market messages was solved, and efficient and accurate identification of black and gray market messages was achieved.

CN122490147APending Publication Date: 2026-07-31CHINA MOBILE FINANCIAL TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE FINANCIAL TECHNOLOGY CO LTD
Filing Date
2026-03-18
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing methods for identifying black and gray market messages are inefficient and inaccurate, and cannot effectively identify hidden illegal messages on the Internet.

Method used

A method combining a large language model and an isolated forest algorithm is adopted. The large language model is trained using a smoothed exponential-logarithmic asymmetric unit activation function, and the isolated forest algorithm is used to process suspicious messages related to black and gray industries to obtain isolated forest scores to determine the identification results.

Benefits of technology

It improves the efficiency of identifying black and gray market messages, reduces the error rate, and can identify black and gray market messages in a timely and effective manner, understand their behavioral trends, and improve the accuracy rate by 30%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122490147A_ABST
    Figure CN122490147A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, storage medium, and program product for identifying black and gray market messages, belonging to the field of artificial intelligence technology. The method includes: acquiring a first message to be identified; inputting the first message into a large language model to obtain suspicious black and gray market messages from the first message output by the large language model; the large language model is trained using a smoothed exponential-logarithmic asymmetric unit activation function and historical black and gray market messages; processing the suspicious black and gray market messages using an isolated forest algorithm to obtain the isolated forest score corresponding to the suspicious black and gray market messages; and determining the suspicious black and gray market messages as black and gray market messages based on the isolated forest score. This solution, based on a large language model and the isolated forest algorithm, identifies black and gray market messages, improving the identification efficiency, reducing the error rate, and enabling timely and effective identification of black and gray market messages, thus understanding the trends of black and gray market activities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and specifically relates to a method, apparatus, device, storage medium and program product for identifying black and gray market information. Background Technology

[0002] With the continuous development of internet technology, fraudulent activities carried out by black and gray industries through the internet are becoming increasingly rampant, making the situation of financial anti-fraud and cybersecurity protection more severe and reaching a stage of intense technical attack and defense.

[0003] Currently, the dissemination of various illegal messages by black and gray market operators to attract traffic online has become a major threat to financial and social security. These illegal messages include information related to financial crimes and other illegal activities published online, severely damaging the security of the financial industry, disrupting financial order, causing significant economic losses to the public, or unknowingly leading people into traps set by criminal groups and making them accomplices. Therefore, it is necessary to identify black and gray market messages online in advance to promptly combat related illegal activities.

[0004] Due to the vast number of websites and the large volume and complexity of messages on the internet, coupled with the relatively covert channels through which black and gray market activities disseminate information online, there are currently no more accurate and efficient technical means to identify such messages. Existing internet black and gray market message identification technologies primarily rely on keyword searches and manual screening, which are not only inefficient but also prone to errors and cannot effectively and promptly identify black and gray market messages. Summary of the Invention

[0005] This application provides a method, apparatus, device, storage medium, and program product for identifying black and gray market information, in order to solve the problems of low efficiency and low accuracy in existing black and gray market information identification methods.

[0006] In a first aspect, embodiments of this application provide a method for identifying black and gray market messages, the method comprising:

[0007] Obtain the first message to be identified;

[0008] The first message is input into a large language model to obtain suspicious black and gray industry messages in the first message output by the large language model; wherein, the large language model is trained using a smoothed exponential-logarithmic asymmetric unit activation function and historical black and gray industry messages;

[0009] The isolated forest algorithm is used to process the suspicious messages related to black and gray industries to obtain the isolated forest score corresponding to the suspicious messages.

[0010] The identification result of the suspicious black and gray industry messages being black and gray industry messages is determined based on the isolated forest score.

[0011] Optionally, the method further includes:

[0012] The smooth exponential-logarithmic asymmetric unit activation function is used as the activation function in the feedforward neural network layer of the large language model.

[0013] Based on the preset number of first word units, the historical black and gray industry messages are converted into first word vectors;

[0014] The first word vector is input into the large language model, and the activation result of the first word vector is determined by the activation function in the feedforward neural network layer.

[0015] The first word vector indicating activation is used to train the feedforward neural network layer to obtain the trained large language model.

[0016] Optionally, the smoothed exponential-logarithmic asymmetric unit activation function is a continuously differentiable smoothing function;

[0017] The smoothed exponential-logarithmic asymmetric unit activation function includes a first segment of differentiable function and a second segment of differentiable function, and the derivatives of the first segment of differentiable function and the second segment of differentiable function are equal.

[0018] Optionally, the step of processing the suspicious messages related to black and gray industries using the isolated forest algorithm to obtain the isolated forest score corresponding to the suspicious messages includes:

[0019] Based on the preset number of second word units, the suspicious messages related to black and gray industries are converted into second word vectors;

[0020] The third word vector is obtained by performing a dot product operation between the preset dimension vector and the second word vector;

[0021] The third word vector is concatenated and aligned to generate a word vector matrix, wherein each row of the word vector matrix includes the third word vector corresponding to the suspicious message from the black and gray market.

[0022] Based on the word vector matrix, the enhancement factor for the isolated forest algorithm is obtained;

[0023] Using the isolated forest algorithm, the isolated forest algorithm enhancement factor, and the average path length of the isolated forest tree corresponding to the suspicious black and gray industry message, the isolated forest score corresponding to the black and gray industry message is obtained.

[0024] Optionally, the word vector matrix includes the third word vectors corresponding to N suspicious messages related to black and gray industries, and the nth row of the word vector matrix includes the third word vector corresponding to the nth suspicious message related to black and gray industries; N is a positive integer, and n is a positive integer greater than or equal to 1 and less than or equal to N;

[0025] The step of obtaining the isolation forest algorithm enhancement factor based on the word vector matrix includes:

[0026] The first difference is obtained based on the m-th third word vector corresponding to the n-th suspected black and gray industry message and the first average value; the first average value is the average value of the m-th third word vector corresponding to each of the N suspected black and gray industry messages; where m is a positive integer greater than or equal to 1 and less than or equal to M, and M is equal to the number of the second word units;

[0027] Obtain the first multiple value between the first difference and the first standard deviation; wherein, the first standard deviation is the standard deviation of the m-th third word vector corresponding to each of the N suspicious black and gray industry messages;

[0028] The weighted average of each first multiple value corresponding to the nth suspicious black and gray industry message is used to obtain the isolation forest algorithm enhancement factor corresponding to the nth suspicious black and gray industry message.

[0029] Optionally, the isolated forest score is between 0 and 1;

[0030] The isolated forest score is directly proportional to the probability that the suspicious black and gray industry information is indeed black and gray industry information.

[0031] Secondly, embodiments of this application also provide a device for identifying black and gray market information, the device comprising:

[0032] The first acquisition module is used to acquire the first message to be identified;

[0033] The first processing module is used to input the first message into a large language model to obtain the suspicious black and gray industry messages in the first message output by the large language model; wherein, the large language model is trained using a smoothed exponential-logarithmic asymmetric unit activation function and historical black and gray industry messages;

[0034] The second processing module is used to process the suspicious messages related to black and gray industries using the isolated forest algorithm to obtain the isolated forest score corresponding to the suspicious messages related to black and gray industries.

[0035] The third processing module is used to determine the identification result of the suspicious black and gray industry messages as black and gray industry messages based on the isolated forest score.

[0036] Thirdly, embodiments of this application also provide a black and gray market message identification device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the black and gray market message identification method as described in any one of the first aspects.

[0037] Fourthly, embodiments of this application also provide a readable storage medium storing a program, which, when executed by a processor, implements the steps of the black and gray market message identification method as described in any one of the first aspects.

[0038] Fifthly, embodiments of this application also provide a computer program product, including computer instructions, which, when executed by a processor, implement the steps in the black and gray market message identification method as described in any one of the first aspects.

[0039] The beneficial effects of this application are:

[0040] The method for identifying black and gray market messages provided in this application involves acquiring a first message to be identified, inputting the first message into a large language model, and obtaining the suspected black and gray market messages in the first message output by the large language model. The large language model is trained using a smoothed exponential-logarithmic asymmetric unit activation function and historical black and gray market messages. The suspected black and gray market messages are processed using an isolated forest algorithm to obtain the corresponding isolated forest score. Based on the isolated forest score, the suspected black and gray market messages are identified as such. Compared to manual identification methods, this method, based on a large language model and the isolated forest algorithm, can improve the identification efficiency of black and gray market messages, reduce the error rate, and effectively and timely identify black and gray market messages, thus understanding the trends of black and gray market activities. Attached Figure Description

[0041] Figure 1 This is a flowchart of the black and gray market message identification method provided in the embodiments of this application;

[0042] Figure 2 This is a graphical schematic diagram of the SELAU activation function provided in the embodiments of this application;

[0043] Figure 3 This is a schematic diagram of the word vector matrix provided in the embodiments of this application;

[0044] Figure 4 This is an overall flowchart of the black and gray market message identification method provided in the embodiments of this application;

[0045] Figure 5 This is a schematic diagram of the structure of the black and gray industry message identification device provided in the embodiments of this application;

[0046] Figure 6 This is a schematic diagram of the structure of the black and gray industry message identification device provided in the embodiments of this application. Detailed Implementation

[0047] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0048] In the following description, specific details such as particular configurations and components are provided merely to aid in a comprehensive understanding of the embodiments of this application. Therefore, those skilled in the art will understand that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Furthermore, for clarity and brevity, descriptions of known functions and constructions have been omitted.

[0049] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0050] In the various embodiments of this application, it should be understood that the sequence number of each process described below does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0051] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and are not used to describe a specified order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, a first object can be one or more. Furthermore, in the specification and claims, "and" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0052] Furthermore, the "or" in this application indicates at least one of the connected objects. For example, "A or B" covers three scenarios: Scenario 1: includes A but excludes B; Scenario 2: includes B but excludes A; Scenario 3: includes both A and B. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0053] The term "instruction" in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as one in which the sender explicitly informs the receiver of specific information, the operation to be performed, or the requested result, etc.; an indirect instruction can be understood as one in which the receiver determines the corresponding information based on the instruction sent by the sender, or makes a judgment and determines the operation to be performed or the requested result, etc., based on the judgment result.

[0054] Before describing the specific embodiments of this application, the following will be explained first:

[0055] In existing technologies for identifying black and gray market messages on the internet, due to the vast number and complexity of websites and messages, and the relatively covert channels through which black and gray market activities disseminate information online, there are no more accurate and efficient technical means for identifying such messages. Currently, the main approach relies on manual retrieval of relevant information from key online channels or automated keyword searches to identify and classify messages related to black and gray market activities. These messages are then manually reviewed, labeled, and categorized before being integrated into a black and gray market intelligence knowledge base. This method primarily depends on the professional experience of human experts and keyword identification of black and gray market messages; it lacks a complete intelligent system, resulting in insufficient accuracy and slow identification speed.

[0056] To address the issues of low efficiency and low accuracy in existing methods for identifying black and gray market messages, this application provides a method, apparatus, device, storage medium, and program product for identifying black and gray market messages.

[0057] like Figure 1 As shown in the figure, this application provides a method for identifying black and gray market messages, the method including:

[0058] Step 101: Obtain the first message to be identified.

[0059] In this step, messages from various websites on the internet are monitored as the first message to be identified. These websites may contain data sources for black and gray market activities to publish related messages. These websites include, but are not limited to, major trading platforms (platforms that serve as channels for black market suppliers to publish and trade), related websites, social media, forums, anonymous group chats, etc.

[0060] Various messages and posts are published on different channels on the Internet. Among them are black and gray market messages (which can also be called illegal messages published by black and gray market operators). By monitoring and viewing the addresses of these posts, a massive source address database of messages published by black and gray market operators can be built.

[0061] We organize and monitor various advertising message sources from black and gray market activities, including the addresses of these sources, and track the aforementioned multiple locations. We receive approximately hundreds of millions of messages daily, allowing us to obtain the first-hand information.

[0062] Step 102: Input the first message into the large language model to obtain the suspicious black and gray industry messages in the first message output by the large language model; wherein, the large language model is trained using a smoothed exponential-logarithmic asymmetric unit activation function and historical black and gray industry messages.

[0063] The Large Language Model (LLM) is trained using a smoothed exponential-logarithmic asymmetric unit activation function and historical black and gray market messages. The trained Large Language Model can also be an improved Large Language Model.

[0064] In this step, the improved Large Language Model (LLM) is used to identify the first message and filter out suspicious messages related to black and gray industries.

[0065] Specifically, the improved Large Language Model (LLM) is used to monitor and retrieve content from internet news sources, identifying suspicious messages from black and gray industries among various online channels.

[0066] Among them, suspicious messages from black and gray industries can also be referred to as suspicious messages released by black and gray industry gangs.

[0067] This step enhances the nonlinear transformation capability of large language models by using asymmetric activation functions.

[0068] Step 103: Use the Enhance Isolation Forest (ELF) algorithm to process the suspicious messages related to black and gray industries, and obtain the corresponding Isolation Forest score value.

[0069] After identifying suspicious messages related to black and gray industries using the improved large language model described above, the isolated forest algorithm is further applied to further identify these messages, obtaining the isolated forest score corresponding to each suspicious message. This isolated forest score is then used to further identify the black and gray industry messages.

[0070] Step 104: Determine the identification result of the suspicious black and gray industry message as a black and gray industry message based on the isolated forest score value.

[0071] In one optional approach, the identification result of black and gray market messages includes whether they are black and gray market messages or not.

[0072] For example, if the isolated forest score is greater than a first preset value, the suspicious black and gray industry message corresponding to the isolated forest score is a black and gray industry message; if the isolated forest score is less than or equal to the first preset value, the suspicious black and gray industry message corresponding to the isolated forest score is not a black and gray industry message.

[0073] In another alternative approach, the identification results of black and gray market messages include the probability that the suspicious black and gray market message is indeed a black and gray market message.

[0074] After obtaining the identification results of black and gray industry messages through this step, black and gray industry messages or messages with a high probability of being black and gray industry messages are identified based on the identification results. The black and gray industry messages or messages with a high probability of being black and gray industry messages are classified and labeled to obtain the classification results. The classification results are then saved and stored in the database.

[0075] The embodiments of this application, through the above steps, identify black and gray industry messages based on large language models and isolated forest algorithms, which can improve the identification efficiency of black and gray industry messages, reduce the identification error rate, and identify black and gray industry messages in a timely and effective manner, so as to understand the trend of black and gray industry behavior.

[0076] In some embodiments, the method further includes:

[0077] The activation function of the smooth exponential logarithmic asymmetric unit (SELAU) is used as the activation function in the feedforward neural network (FFN or FNN) layer of the large language model.

[0078] It should be noted that the feedforward neural network layer is an important component of the Large Language Model (LLM). The activation function is the core of the neuron computation in this layer. The activation function makes the input and output of the Large Language Model (LLM) no longer linear. Since there is no linear input and output relationship between language conversation text, the existence of the activation function can better fit the input and output function relationship that is closer to reality.

[0079] The original activation function in the large language model is the Rectified Linear Unit (ReLU) activation function.

[0080] It is understandable that the application adapter module method is used to improve the large language model LLM by releasing the parameters of the feedforward neural network layer (FNN) and locking the parameters of other models, mainly by replacing the activation function with the SELAU activation function.

[0081] Smoothed log-asymmetric unit (SELAU) is a newly created activation function used in the fully connected feedforward neural network of large language models (LLMs). SELAU not only possesses the properties of a nonlinear continuous function mapping but also prevents gradient explosion and maintains sparsity by not activating random neurons. This saves computational resources, prevents overfitting, and compensates for the original ReLU activation function's problem of setting all negative inputs to zero, enhancing its ability to learn from negative inputs. Overall, the SELAU activation function retains the advantages of the ReLU activation function while eliminating its disadvantages such as potential gradient explosion, setting negative inputs to zero, and lack of smoothness, thus improving the training and recognition efficiency of large language models (LLMs) for unstructured message data.

[0082] Based on the preset number of first word units, the historical black and gray industry messages are converted into first word vectors.

[0083] It should be noted that during the training (or fine-tuning) of the large language model LLM, samples of relevant black market transaction messages released by black market gangs will be input, i.e., historical black market messages.

[0084] First, prepare historical black and gray market messages, such as: text1: "Selling [various types of mobile phone cards] In stock, various phone cards, regional phone cards, registered cards, real-name phone cards, real cards with SMS verification code, after-sales guarantee. Sincere recruitment of agents. For advertising accounts, please contact TG: @*********. Two-way contact: @*********."

[0085] As mentioned above, in historical black and gray market messages, each word is a token that can be converted into a word vector.

[0086] There are approximately 300 million historical black and gray market messages. The maximum number of word units in each message is extracted and converted into word vectors (i.e., the first word vectors), which are then input into the improved Large Language Model (LLM). That is, the number of first word units is 500.

[0087] The first word vector is input into the large language model, and the activation result of the first word vector is determined by the activation function in the feedforward neural network layer.

[0088] Specifically, the first word vector is input into the feedforward neural network layer after matrix operations, and the activation result of the first word vector is determined by the SELAU activation function of the smoothed exponential-logarithmic asymmetric unit in the feedforward neural network layer.

[0089] The first word vector indicating activation is used to train the feedforward neural network layer to obtain the trained large language model.

[0090] The neurons in the feedforward neural network layer activate the first word vectors through the SELAU activation function, and the weighted output is calculated. The output is then applied to backpropagation to update the parameter matrix of the feedforward neural network layer. In other words, the two fully connected layers in the large language model are used for parameter fine-tuning training, serving as a dimensionality reduction layer and a dimensionality increase layer.

[0091] Subsequently, the improved large language model LLM is prompted and fine-tuned to guide the output of correct recognition results.

[0092] For example, a prompt template: Please identify whether the message in the network address is a black and gray market message: {text1}, please identify if it is and save the record of this message.

[0093] Optionally, the smoothed exponential-logarithmic asymmetric unit activation function is a continuously differentiable smoothing function;

[0094] The smoothed exponential-logarithmic asymmetric unit activation function includes a first segment of differentiable function and a second segment of differentiable function, and the derivatives of the first segment of differentiable function and the second segment of differentiable function are equal.

[0095] This embodiment applies an improved Large Language Model (LLM) technique to monitor and retrieve content from internet news sources, identifying information related to black and gray market activities across various online channels. It primarily utilizes the properties of asymmetric activation functions to enhance the nonlinear transformation capabilities of the large model, as detailed below:

[0096] The SELAU activation function provided in this embodiment is a piecewise calculation function, which can be applied to fitting functions for different input and output scenarios. The specific formula is as follows:

[0097]

[0098] Where α, β, and γ are natural numbers greater than zero, generally within the range (0, 10]. Optionally, α = 3, β = 1, and γ = 1. e is a natural constant, approximately equal to 2.718282 (rounded to six decimal places). x represents the weighted value of the word vector converted from the token in the message sentence after matrix transformation.

[0099] In the above formula for calculating the SELAU activation function, in Time is linear Functions and exponents Multiplication, in Logarithm The functions are both piecewise functions, each passing through the origin. Taking the derivatives of the first and second piecewise functions respectively, the results are as follows:

[0100] Differentiate the first segment of the differentiable function: ;

[0101] Differentiate the second differentiable function: ;

[0102] The above piecewise differentiable function at the origin The derivatives after substituting the expression at =0 into the formula are all... Set during training The constant term is 3. If the derivative is 1, then the derivative of both differentiable functions is 3, which indicates that the SELAU activation function is a continuously differentiable smooth function. The specific function graph is shown below. Figure 2 As shown.

[0103] Compared to the default ReLU activation function in the original large language model LLM, the SELAU activation function can improve the recognition efficiency in scenarios such as the posting of black and gray market transaction messages during training and recognition. Identifying black and gray market messages from massive amounts of internet data requires a smooth, non-linear activation function for mapping, which can better learn the information differences between different message contexts. Simultaneously, during training, setting the activation function to 0 for large negative inputs keeps neurons inactive, maintaining sparsity and increasing computational efficiency. It also ensures that some negative inputs produce negative outputs, avoiding information loss. For positive inputs, the slow growth of the logarithmic function prevents excessively large activations that could cause gradient explosion.

[0104] The SELAU activation function combines the characteristics of the ReLU activation function with the disadvantages of non-smoothness and the loss of information caused by negative inputs (all zeros). The large model improved by the SELAU activation function improves the efficiency of identifying black and gray market messages and enhances the ability of the large model to distinguish black and gray market messages. After applying the SELAU activation function, the accuracy of black and gray market message identification increased by 30%.

[0105] After training to obtain the large language model, the first message is input into the large oracle model to obtain the suspicious black and gray industry messages in the first message output by the large language model.

[0106] Optionally, the first message is converted into prompt words. Then, the finely tuned Large Language Model (LLM) described above is used to simulate the recognition task through prompt words to identify suspicious messages from the black and gray industries on the Internet. Various prompt words such as Internet message source links are loaded into the dialogue to simulate the recognition task and then the suspicious messages from the black and gray industries are identified.

[0107] For example, if we need to identify and extract the first message from a website, the generated prompt words would be as follows:

[0108] Character Description:

[0109] "You are a black and gray market content identification expert. Your task is to find suspicious black and gray market messages published by black and gray market operators from the provided website links and return the content results of the suspicious black and gray market messages."

[0110] Simulated task dialogue input and output:

[0111] Input: The content of the first message {"Selling and buying various cards... For advertising accounts, please contact: @******"}. When encountering the above first message content, it is judged as suspicious black and gray industry content. Please completely identify the original first message content and extract and record it in the specified location "******".

[0112] Output: Okay, the content of the aforementioned suspicious messages from the black and gray market has been extracted and recorded in its original form.

[0113] Task Description:

[0114] You need to go to this network address "https: / / ***********" to view and search all the published content messages, find the suspicious black and gray market messages, and extract all the records to the specified location "******".

[0115] This embodiment improves the nonlinear mapping capability of the feedforward neural network layer of a large language model by using the SELAU activation function, thereby enhancing the efficiency of the large language model in identifying suspicious messages from black and gray market activities. Furthermore, it combines this with a prompt word simulation task process to improve the accuracy of the large language model's recognition. Therefore, by applying the SELAU activation function to improve the large language model and combining it with a prompt word simulation task process to identify suspicious messages from black and gray market activities on the internet and extract records, the recognition accuracy reached over 95%, efficiency was improved by about 20%, and overall accuracy was improved by about 30%.

[0116] In some embodiments, the process of using the isolated forest algorithm to process the suspicious messages related to black and gray industries to obtain the isolated forest score corresponding to the suspicious messages includes:

[0117] Based on the preset number of second word units, the suspicious messages related to black and gray industries are converted into second word vectors.

[0118] Optionally, the number of suspicious messages related to black and gray industries is 10,000, and the number of second word units is 120 word units.

[0119] Specifically, a suspicious message related to black and gray market activities is a sample message. For example, "Selling [various mobile phone SIM cards] phone cards, virtual cards, phone cards from various countries..." is a sample message. The sample message includes word units with the second word unit count, and "phone cards, virtual cards, phone cards from various countries..." in the sample message is the second word vector.

[0120] The third word vector is obtained by performing a dot product operation between the preset dimension vector and the second word vector.

[0121] Specifically, in this embodiment, a word vector normalization alignment method is used to enable unstructured data to be used for training in the unsupervised isolated forest algorithm. The calculation formula is as follows: Here It is a 1×512 dimension constant vector (i.e., a preset dimension vector) and Transposing the vector and performing a dot product operation yields the normalized word vector. ;in, This represents the second word vector. This represents the third word vector.

[0122] In this embodiment, the suspicious messages related to black and gray market activities consist of 10,000 sample message points. Each sample message point is generally composed of word units (tokens) ranging from 100 to 500. Each sample message point is truncated to 120 word units, and the corresponding large model word vector for each word unit is... , It consists of 1×512 dimensions.

[0123] The third word vector is concatenated and aligned to generate a word vector matrix, wherein each row of the word vector matrix includes the third word vector corresponding to a suspicious message from the black and gray industries.

[0124] This involves concatenating the aligned positions of multiple sample points to obtain a word vector matrix.

[0125] For example, the first column of the word vector matrix contains N sample messages (i.e., N suspicious messages related to black and gray industries), and each row of the word vector matrix includes the third word vector corresponding to the sample message. The word vector matrix is ​​as follows: Figure 3 As shown. The word vector matrix includes the third word vectors corresponding to N suspicious messages (sample messages) related to black and gray industries. The nth row of the word vector matrix includes the third word vector corresponding to the nth suspicious message related to black and gray industries; N is a positive integer, and n is a positive integer greater than or equal to 1 and less than or equal to N. Where m is a positive integer greater than or equal to 1 and less than or equal to M, and M is equal to the number of second word units.

[0126] Based on the word vector matrix, the enhancement factor of the isolated forest algorithm is obtained. .

[0127] Using the Isolation Forest algorithm, the Isolation Forest algorithm enhancement factor, and the average path length of the isolated forest trees corresponding to the suspicious black and gray industry messages, the Isolation Forest score corresponding to the black and gray industry messages is obtained. .

[0128] The step of obtaining the isolation forest algorithm enhancement factor based on the word vector matrix includes:

[0129] A first difference is obtained based on the m-th third word vector corresponding to the n-th suspected black and gray industry message and a first average value; the first average value is the average value of the m-th third word vector corresponding to each of the N suspected black and gray industry messages; where m is a positive integer greater than or equal to 1 and less than or equal to M, and M is equal to the number of the second word units. Optionally, N equals 10000.

[0130] Obtain the first multiple value between the first difference and the first standard deviation; wherein, the first standard deviation is the standard deviation of the m-th third word vector corresponding to each of the N suspicious black and gray industry messages.

[0131] The weighted average of each first multiple value corresponding to the nth suspicious black and gray industry message is used to obtain the isolation forest algorithm enhancement factor corresponding to the nth suspicious black and gray industry message.

[0132] The formula for calculating the isolated forest score corresponding to the nth suspicious message from the black and gray market is as follows:

[0133]

[0134] Specifically:

[0135]

[0136]

[0137]

[0138]

[0139] in, The enhancement factor for the Isolation Forest algorithm is the third word vector in each sample message point. The average of the third word vectors in all corresponding sample message points The difference is then divided by the standard deviation of all corresponding sample message points. Calculate the multiple relationship, and finally take a weighted average of all third-word vectors for each sample message point. It represents the number of third word vectors in each sample message point. This represents the average path length of the isolated forest tree corresponding to suspicious messages from black and gray industries. It is calculated using a fixed method and is related to the number of sample messages, N. This represents the average path length of all isolated forest trees across all sample message points. Indicates sample message point The number of edges traversed from the root node to a leaf node in each isolated forest tree is called the path length, and 𝑘 is the number of isolated trees.

[0140] Based on the improved large language model described above, suspicious black and gray market messages are identified and extracted from network messaging platforms. Further unsupervised learning techniques, including the Enhanced Isolation Forest algorithm and word vector normalization alignment, are applied to detect anomalies in these messages. Detected anomalies are marked as high-priority suspicious messages and saved to a database. The Enhanced Isolation Forest algorithm (EIF) is primarily used to detect anomalies in unstructured message content.

[0141] In some embodiments, the isolated forest score is between 0 and 1;

[0142] The isolated forest score is directly proportional to the probability that the suspicious black and gray industry information is indeed black and gray industry information.

[0143] Understandably, the closer the Isolation Forest Score of a suspicious message related to black and gray market activities is to 1, the greater the likelihood that the message is indeed related to black and gray market activities. Specifically, when the Isolation Forest Score is greater than 0.5 and less than or equal to 1 (i.e., the Isolation Forest Score falls within the range of (0.5, 1)), the probability that the message is indeed related to black and gray market activities is high (i.e., there is an anomaly). Conversely, when the Isolation Forest Score is greater than 0 and less than or equal to 0.5 (i.e., the Isolation Forest Score falls within the range of (0, 0.5)), the probability that the message is indeed related to black and gray market activities is low.

[0144] Optionally, the score threshold can be set to 0.8.

[0145] If the isolated forest score is greater than the score threshold (i.e., the isolated forest score is within the range of [0.8,1]), suspicious black and gray industry messages are marked as "yes" (indicating that the suspicious black and gray industry message is highly likely to be a black and gray industry message or can be considered to be a black and gray industry message), and need to be given special attention; otherwise, they are marked as "no" (indicating that the suspicious black and gray industry message is not highly likely to be a black and gray industry message or is not a black and gray industry message).

[0146] This embodiment uses a weighted average method, comparing the difference between word vectors and their average value to a multiple of the word vector standard deviation, to calculate the θ value for each sample message point. This serves as an enhancement factor for the Isolation Forest algorithm, accelerating its convergence speed and improving its efficiency in anomaly detection. It can more quickly detect special anomalies in black and gray market messages. Simultaneously, it applies a word vector normalization alignment method, enabling the unsupervised Isolation Forest algorithm to use unstructured data. Typically, structured numerical data is input into the Isolation Forest algorithm for computation; this embodiment makes it possible to use unstructured text message data for model computation. This embodiment demonstrates that the enhanced Isolation Forest algorithm improves efficiency by approximately 35% compared to the original Isolation Forest algorithm in detecting special anomalies in black and gray market messages.

[0147] Optionally, suspicious messages marked as "yes" and suspicious messages marked as "no" can be sorted and recorded in the database. For example, the data format of the sorted records is shown in Table 1 below.

[0148] Table 1

[0149]

[0150] Based on the black and gray market messages recorded after anomaly detection, we pay special attention to messages marked as "yes" as suspicious, as these may be black and gray market messages. We can provide clues to take action against them or play a role in risk prevention and control for related businesses, thereby improving our ability to prevent and control risks from the black and gray market.

[0151] The following is combined Figure 4 The following describes the overall process of the black and gray market message identification method provided in this embodiment:

[0152] Obtain first messages from internet channel message data sources; use a large language model to identify suspicious black and gray industry messages in the first messages; detect and label whether the suspicious black and gray industry messages are indeed black and gray industry messages; organize and store the labeled suspicious black and gray industry messages in the database.

[0153] In summary, the black and gray market message identification method provided in this application is based on an asymmetric function large model and unsupervised learning techniques. It applies an improved Large Language Model (LLM) technique combined with random forest unsupervised learning, enabling unsupervised learning algorithms to be applied to unstructured data detection. Simultaneously, the SELAU activation function is used to improve the large model's identification efficiency, accurately identifying new messages from black and gray market activities from massive amounts of text, social media, forums, and other channels on various websites and websites. These identified messages are then integrated and stored in a database. In other words, this application proposes a black and gray market message identification method based on an improved asymmetric large model and enhanced unsupervised techniques, improving the efficiency of black and gray market message identification. It also proposes an anomaly detection method for black and gray market messages using an unsupervised isolated forest algorithm on unstructured text data, making unsupervised learning a feasible approach for training unstructured text message data.

[0154] By monitoring massive amounts of messages from various internet platforms and websites, an improved Large Language Model (LLM) technique is used to more accurately and effectively identify and filter suspicious messages published by black and gray market actors. This is then combined with an unsupervised learning-based reinforced isolated forest algorithm to detect related messages from black and gray market actors, identifying and labeling special abnormal messages as a set of messages to focus on. Finally, the identified and labeled black and gray market messages are organized and stored in a database for various organizations to access and view, enabling them to perceive whether their own businesses have vulnerabilities and risks, and to take timely risk prevention and control measures.

[0155] This application proposes a Smoothed Logarithmic Asymmetric Unit (SELAU) activation function, which is loaded into the feedforward neural network layer of a large language model (LLM) for nonlinear mapping. Utilizing the smooth, continuous, and asymmetric properties of this activation function improves the large model's ability to identify black and gray market information. It also proposes an enhanced isolated forest algorithm, using the relationship between data points and the mean difference relative to the standard deviation to establish an enhancement factor, thereby improving the anomaly detection efficiency of the original isolated forest algorithm and enabling faster identification of special abnormal messages in black and gray market messages. Furthermore, it proposes a word vector normalization alignment method, which can apply unstructured text message content data to unsupervised learning, making anomaly detection of unstructured data using the isolated forest algorithm a feasible approach. Finally, it proposes a method for identifying and extracting records of black and gray market messages using asymmetric improvements to the large model and enhanced unsupervised techniques.

[0156] Specifically, based on the existing black and gray market message dissemination business environment on various internet platforms, a Smoothed Exponential-Logarithmic Asymmetric Unit (SELAU) activation function is created to improve the original large language model feedforward neural network layer. This makes the nonlinear mapping of the large model feedforward neural network layer smoother, retains the advantages of the original ReLU activation function, effectively prevents gradient explosion, and can also preserve negative input information. It has a better learning and recognition ability for massive amounts of black and gray market message content on the internet, improves the recognition efficiency of black and gray market transaction messages, and saves a lot of time compared to manually searching and collecting black and gray market messages from massive amounts of internet information. The data standard deviation multiple method is used to construct enhancement factors as factor coefficients for the isolated forest algorithm, improving the anomaly detection efficiency of the original isolated forest algorithm. At the same time, a word vector normalization alignment method is designed to make it possible to train the unsupervised isolated forest algorithm for anomaly detection using unstructured text content data input. Currently, manually loading some programs to identify messages published by black and gray market entities on the Internet is prone to errors and is inefficient. This paper proposes a feasible method for identifying and extracting records of black and gray market information on the Internet. Compared with manual retrieval and identification, it not only saves labor costs, but also improves the efficiency and accuracy of record identification and extraction.

[0157] like Figure 5 As shown in the figure, this application embodiment also provides a black and gray market message identification device, the device comprising:

[0158] The first acquisition module 501 is used to acquire the first message to be identified;

[0159] The first processing module 502 is used to input the first message into a large language model to obtain the suspicious black and gray industry messages in the first message output by the large language model; wherein, the large language model is trained using a smoothed exponential-logarithmic asymmetric unit activation function and historical black and gray industry messages;

[0160] The second processing module 503 is used to process the suspicious messages related to black and gray industries using the isolated forest algorithm to obtain the isolated forest score corresponding to the suspicious messages related to black and gray industries.

[0161] The third processing module 504 is used to determine the identification result of the suspicious black and gray industry message as a black and gray industry message based on the isolated forest score value.

[0162] Optionally, the device further includes:

[0163] The fourth processing module is used to use the smoothed exponential-logarithmic asymmetric unit activation function as the activation function in the feedforward neural network layer of the large language model;

[0164] The fifth processing module is used to convert the historical black and gray industry messages into first word vectors according to the preset number of first word units;

[0165] The sixth processing module is used to input the first word vector into the large language model and use the activation function in the feedforward neural network layer to determine the activation result of the first word vector;

[0166] The seventh processing module is used to train the feedforward neural network layer using the activation result indicating the activation of the first word vector, so as to obtain the trained large language model.

[0167] Optionally, the smoothed exponential-logarithmic asymmetric unit activation function is a continuously differentiable smoothing function;

[0168] The smoothed exponential-logarithmic asymmetric unit activation function includes a first segment of differentiable function and a second segment of differentiable function, and the derivatives of the first segment of differentiable function and the second segment of differentiable function are equal.

[0169] Optionally, the second processing module 503 includes:

[0170] The first processing unit is used to convert the suspicious black and gray industry messages into second word vectors according to a preset number of second word units;

[0171] The second processing unit is used to perform a dot product operation between the preset dimension vector and the second word vector to obtain the third word vector;

[0172] The third processing unit is used to perform concatenation and alignment processing on the third word vector to generate a word vector matrix, wherein each row of the word vector matrix includes a third word vector corresponding to the suspicious message from the black and gray industries.

[0173] The fourth processing unit is used to obtain the isolation forest algorithm enhancement factor based on the word vector matrix;

[0174] The fifth processing unit is used to obtain the isolated forest score corresponding to the black and gray industry message by using the isolated forest algorithm, the isolated forest algorithm enhancement factor, and the average path length of the isolated forest tree corresponding to the suspicious message.

[0175] Optionally, the word vector matrix includes the third word vectors corresponding to N suspicious messages related to black and gray industries, and the nth row of the word vector matrix includes the third word vector corresponding to the nth suspicious message related to black and gray industries; N is a positive integer, and n is a positive integer greater than or equal to 1 and less than or equal to N;

[0176] The fourth processing unit is specifically used for:

[0177] The first difference is obtained based on the m-th third word vector corresponding to the n-th suspected black and gray industry message and the first average value; the first average value is the average value of the m-th third word vector corresponding to each of the N suspected black and gray industry messages; where m is a positive integer greater than or equal to 1 and less than or equal to M, and M is equal to the number of the second word units;

[0178] Obtain the first multiple value between the first difference and the first standard deviation; wherein, the first standard deviation is the standard deviation of the m-th third word vector corresponding to each of the N suspicious black and gray industry messages;

[0179] The weighted average of each first multiple value corresponding to the nth suspicious black and gray industry message is used to obtain the isolation forest algorithm enhancement factor corresponding to the nth suspicious black and gray industry message.

[0180] Optionally, the isolated forest score is between 0 and 1;

[0181] The isolated forest score is directly proportional to the probability that the suspicious black and gray industry information is indeed black and gray industry information.

[0182] It should be noted that the black and gray industry message identification device provided in this application embodiment is a device capable of executing the above-described black and gray industry message identification method. Therefore, all embodiments of the above-described black and gray industry message identification method are applicable to this device and can achieve the same or similar technical effects.

[0183] like Figure 6 As shown in the figure, this application embodiment also provides a black and gray industry message identification device, including: a processor 601; and a memory 603 connected to the processor 601 via a bus interface 602. The memory 603 is used to store the programs and data used by the processor 601 when performing operations, and the processor 601 calls and executes the programs and data stored in the memory 603.

[0184] The transceiver 604 is connected to the bus interface 602 and is used to receive and send data under the control of the processor 601. Specifically, the processor 601 is used to read the program in the memory 603 and to execute the following processes:

[0185] Obtain the first message to be identified;

[0186] The first message is input into a large language model to obtain suspicious black and gray industry messages in the first message output by the large language model; wherein, the large language model is trained using a smoothed exponential-logarithmic asymmetric unit activation function and historical black and gray industry messages;

[0187] The isolated forest algorithm is used to process the suspicious messages related to black and gray industries to obtain the isolated forest score corresponding to the suspicious messages.

[0188] The identification result of the suspicious black and gray industry messages being black and gray industry messages is determined based on the isolated forest score.

[0189] Optionally, the processor 601 is further configured to:

[0190] The smooth exponential-logarithmic asymmetric unit activation function is used as the activation function in the feedforward neural network layer of the large language model.

[0191] Based on the preset number of first word units, the historical black and gray industry messages are converted into first word vectors;

[0192] The first word vector is input into the large language model, and the activation result of the first word vector is determined by the activation function in the feedforward neural network layer.

[0193] The first word vector indicating activation is used to train the feedforward neural network layer to obtain the trained large language model.

[0194] Optionally, the smoothed exponential-logarithmic asymmetric unit activation function is a continuously differentiable smoothing function;

[0195] The smoothed exponential-logarithmic asymmetric unit activation function includes a first segment of differentiable function and a second segment of differentiable function, and the derivatives of the first segment of differentiable function and the second segment of differentiable function are equal.

[0196] Optionally, the processor 601 is used for:

[0197] Based on the preset number of second word units, the suspicious messages related to black and gray industries are converted into second word vectors;

[0198] The third word vector is obtained by performing a dot product operation between the preset dimension vector and the second word vector;

[0199] The third word vector is concatenated and aligned to generate a word vector matrix, wherein each row of the word vector matrix includes the third word vector corresponding to the suspicious message from the black and gray market.

[0200] Based on the word vector matrix, the enhancement factor for the isolated forest algorithm is obtained;

[0201] Using the isolated forest algorithm, the isolated forest algorithm enhancement factor, and the average path length of the isolated forest tree corresponding to the suspicious black and gray industry message, the isolated forest score corresponding to the black and gray industry message is obtained.

[0202] Optionally, the word vector matrix includes the third word vectors corresponding to N suspicious messages related to black and gray industries, and the nth row of the word vector matrix includes the third word vector corresponding to the nth suspicious message related to black and gray industries; N is a positive integer, and n is a positive integer greater than or equal to 1 and less than or equal to N;

[0203] The processor 601 is specifically used for:

[0204] The first difference is obtained based on the m-th third word vector corresponding to the n-th suspected black and gray industry message and the first average value; the first average value is the average value of the m-th third word vector corresponding to each of the N suspected black and gray industry messages; where m is a positive integer greater than or equal to 1 and less than or equal to M, and M is equal to the number of the second word units;

[0205] Obtain the first multiple value between the first difference and the first standard deviation; wherein, the first standard deviation is the standard deviation of the m-th third word vector corresponding to each of the N suspicious black and gray industry messages;

[0206] The weighted average of each first multiple value corresponding to the nth suspicious black and gray industry message is used to obtain the isolation forest algorithm enhancement factor corresponding to the nth suspicious black and gray industry message.

[0207] Optionally, the isolated forest score is between 0 and 1;

[0208] The isolated forest score is directly proportional to the probability that the suspicious black and gray industry information is indeed black and gray industry information.

[0209] Among them, Figure 6 In this context, the bus architecture may include any number of interconnected buses and bridges, specifically linking various circuits together, represented by one or more processors (processor 601) and memory (memory 603). The bus architecture may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. A bus interface provides a user interface 605. A transceiver 604 may be multiple elements, including transmitters and receivers, providing units for communicating with various other devices over a transmission medium. Processor 601 is responsible for managing the bus architecture and general processing, and memory 603 may store data used by processor 601 during operation.

[0210] In addition, specific embodiments of this application also provide a readable storage medium storing a computer program thereon, wherein when the program is executed by a processor, it implements the steps in the black and gray market message identification method as described above.

[0211] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0212] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can be physically included separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0213] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions that cause a computer device (which may be a personal computer, server, or network device, etc.) to execute partial steps of the resource selection method described in the various embodiments of this application, or to execute partial steps of the information transmission method described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0214] A specific embodiment of this application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above-described functionality. Figure 1 The various processes of the method embodiments shown can achieve the same technical effect, and will not be described again here to avoid repetition.

[0215] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for identifying black hat production messages, characterized by, The method includes: Obtain the first message to be identified; The first message is input into a large language model to obtain suspicious black and gray industry messages in the first message output by the large language model; wherein, the large language model is trained using a smoothed exponential-logarithmic asymmetric unit activation function and historical black and gray industry messages; The isolated forest algorithm is used to process the suspicious messages related to black and gray industries to obtain the isolated forest score corresponding to the suspicious messages. The identification result of the suspicious black and gray industry messages being black and gray industry messages is determined based on the isolated forest score.

2. The method of claim 1, wherein, The method further includes: The smooth exponential-logarithmic asymmetric unit activation function is used as the activation function in the feedforward neural network layer of the large language model. Based on the preset number of first word units, the historical black and gray industry messages are converted into first word vectors; The first word vector is input into the large language model, and the activation result of the first word vector is determined by the activation function in the feedforward neural network layer. The first word vector indicating activation is used to train the feedforward neural network layer to obtain the trained large language model.

3. The method according to claim 1 or 2, characterized in that, The smoothed exponential-logarithmic asymmetric unit activation function is a continuously differentiable smoothing function; The smoothed exponential-logarithmic asymmetric unit activation function includes a first segment of differentiable function and a second segment of differentiable function, and the derivatives of the first segment of differentiable function and the second segment of differentiable function are equal.

4. The method of claim 1, wherein, The process of using the isolated forest algorithm to process the suspicious messages related to black and gray industries, and obtaining the isolated forest score corresponding to the suspicious messages, includes: Based on the preset number of second word units, the suspicious messages related to black and gray industries are converted into second word vectors; The third word vector is obtained by performing a dot product operation between the preset dimension vector and the second word vector; The third word vector is concatenated and aligned to generate a word vector matrix, wherein each row of the word vector matrix includes the third word vector corresponding to the suspicious message from the black and gray market. Based on the word vector matrix, the enhancement factor for the isolated forest algorithm is obtained; Using the isolated forest algorithm, the isolated forest algorithm enhancement factor, and the average path length of the isolated forest tree corresponding to the suspicious black and gray industry message, the isolated forest score corresponding to the black and gray industry message is obtained.

5. The method of claim 4, wherein, The word vector matrix includes the third word vectors corresponding to N suspicious messages from black and gray industries, and the nth row of the word vector matrix includes the third word vector corresponding to the nth suspicious message from black and gray industries; N is a positive integer, and n is a positive integer greater than or equal to 1 and less than or equal to N; The step of obtaining the isolation forest algorithm enhancement factor based on the word vector matrix includes: The first difference is obtained based on the m-th third word vector corresponding to the n-th suspected black and gray industry message and the first average value; the first average value is the average value of the m-th third word vector corresponding to each of the N suspected black and gray industry messages; where m is a positive integer greater than or equal to 1 and less than or equal to M, and M is equal to the number of the second word units; Obtain the first multiple value between the first difference and the first standard deviation; wherein, the first standard deviation is the standard deviation of the m-th third word vector corresponding to each of the N suspicious black and gray industry messages; The weighted average of each first multiple value corresponding to the nth suspicious black and gray industry message is used to obtain the isolation forest algorithm enhancement factor corresponding to the nth suspicious black and gray industry message.

6. The method of claim 1, wherein, The isolated forest score is between 0 and 1; The isolated forest score is directly proportional to the probability that the suspicious black and gray industry information is indeed black and gray industry information.

7. A black hat production message identification apparatus characterized by comprising: The device includes: The first acquisition module is used to acquire the first message to be identified; The first processing module is used to input the first message into a large language model to obtain the suspicious black and gray industry messages in the first message output by the large language model; wherein, the large language model is trained using a smoothed exponential-logarithmic asymmetric unit activation function and historical black and gray industry messages; The second processing module is used to process the suspicious messages related to black and gray industries using the isolated forest algorithm to obtain the isolated forest score corresponding to the suspicious messages related to black and gray industries. The third processing module is used to determine the identification result of the suspicious black and gray industry messages as black and gray industry messages based on the isolated forest score.

8. A black hat production message identification device characterized by comprising: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the black and gray market message identification method as described in any one of claims 1 to 6.

9. A readable storage medium, characterized in that, The readable storage medium stores a program that, when executed by a processor, implements the steps of the black and gray market message identification method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps in the black and gray market message identification method as described in any one of claims 1 to 6.