Information identification methods, devices, and electronic devices

The method and device leverage pre-trained models and keyword/regular expression libraries to identify abnormal gaming accounts and user identities, addressing privacy constraints and enhancing gaming safety and user experience.

JP7839834B2Active Publication Date: 2026-04-02NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

In overseas gaming markets, the protection of user privacy and personal information is stringent, making it difficult for game companies to identify and prevent illegal activities induced by unscrupulous individuals, especially when minors are involved, as they cannot obtain identity information or investigate user statements, leading to credibility issues and user loss.

Method used

An information identification method and device that utilizes behavior data and speech content to perform risk assessment and anomaly identification, employing pre-trained models to analyze game account data for abnormal content, probability of abnormal accounts, and potential user identity, integrating regular expressions and keyword libraries for rapid detection.

Benefits of technology

Provides a safe gaming environment by accurately identifying abnormal accounts and user identities, enhancing user experience through efficient detection and prevention of illegal activities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007839834000001
    Figure 0007839834000001
  • Figure 0007839834000002
    Figure 0007839834000002
  • Figure 0007839834000003
    Figure 0007839834000003
Patent Text Reader

Abstract

To provide an information identification method, a device, and an electronic device for improving a user's game experience by providing the user with a safe game environment by determining an abnormal game account and remark in a game.SOLUTION: A method acquires an identification result by acquiring first remark data corresponding to a first game account, acquires account information of the first game account when it is shown that the remark data include abnormal contents, stores it in a preset database in association with the first remark data, inputs account information of the first game account to a first model, acquires a possibility of an abnormal account, inputs the first remark data in a preset database to a second model when the probability is larger than a preset probability threshold, and acquires at least one piece of information from among normality of a remark, abnormality information corresponding to the first remark data, and contact information and identity information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and particularly to an information identification method, apparatus, and electronic device.

Background Art

[0002] In some popular games, there is a phenomenon that some illegal actors induce minors to a third-party platform to commit illegal acts. This phenomenon has an adverse impact on the game, seriously damages the credibility of the game, and leads to the loss of users. However, in the overseas market, there are strict requirements for the protection of privacy and personal information. Overseas game companies cannot obtain the identity information of users and do not have the right to arbitrarily investigate the speech records of users. Next, illegal actors often induce minors to a third-party platform, and it is difficult to leave illegal evidence, and game companies cannot obtain sufficient clues.

Summary of the Invention

[0003] An object of the present invention is to provide an information identification method, apparatus, and electronic device that can use the behavior data and speech content of users to perform risk assessment and anomaly identification, provide a safe game environment for users, and improve the game experience of users.

[0004] According to a first aspect, the present invention provides an information identification method, which includes: acquiring first statement data corresponding to a first game account in a target game; performing abnormal content identification processing on the first statement data and obtaining an identification result; if the identification result indicates that the first statement data contains abnormal content, storing the first statement data in a pre-configured database; acquiring account information of the first game account; associating the account information of the first game account with the first statement data and storing it in a pre-configured database; inputting the account information of the first game account in the pre-configured database into a pre-trained first model and obtaining a first output result, wherein the first output result is used to indicate the probability that the first game account is an abnormal account; if the probability corresponding to the first output result is greater than a pre-configured probability threshold, inputting the first statement data in the pre-configured database into a pre-trained second model and obtaining a second output result, wherein the second output result includes at least one of the following: that the statement is normal, abnormal information corresponding to the first statement data, contact information, and user identity information.

[0005] According to a second aspect, the present invention provides an information identification device comprising a data acquisition module, an information storage module, a first identification module, and a second identification module, wherein the data acquisition module is configured to acquire first statement data corresponding to a first game account in a target game, perform abnormal content identification processing on the first statement data, and acquire the identification result, and the information storage module is configured to store the first statement data in a pre-configured database if the identification result indicates that the first statement data contains abnormal content, acquire account information of the first game account, and store the account information of the first game account and the first statement data in a pre-configured database in association with each other. The first identification module is configured to input account information of a first game account from a pre-configured database into a pre-trained first model and obtain a first output result, which is used to indicate the probability that the first game account is an abnormal account. The second identification module is configured to input first utterance data from a pre-configured database into a pre-trained second model and obtain a second output result, which, if the probability corresponding to the first output result is greater than a pre-configured probability threshold, includes at least one of the following: that the utterance is normal, abnormal information corresponding to the first utterance data, contact information, and user identity information.

[0006] According to a third aspect, the present invention provides an electronic device including a processor and a memory, wherein the memory stores device-executable instructions that can be executed by the processor, and the processor executes the device-executable instructions to realize the information identification method.

[0007] According to a fourth aspect, the present invention provides a computer-readable storage medium in which computer executable instructions are stored, and when the computer executable instructions are called and executed by a processor, the computer executable instructions cause the processor to implement the information identification method described above.

[0008] The embodiments of this invention provide the following beneficial effects.

[0009] The information identification method, apparatus, and electronic device according to the present invention first acquire first statement data corresponding to a first game account in a target game, perform abnormal content identification processing on the first statement data, obtain an identification result, and if the identification result indicates that the first statement data contains abnormal content, store the first statement data in a pre-configured database, acquire account information of the first game account, associate the account information of the first game account with the first statement data and store it in a pre-configured database, then input the account information of the first game account in the pre-configured database into a pre-trained first model and obtain a first output result, which is used to indicate the probability that the first game account is an abnormal account, and if the probability corresponding to the first output result is greater than a pre-configured probability threshold, input the first statement data in the pre-configured database into a pre-trained second model and obtain a second output result, which includes at least one of the following: that the statement is normal, abnormal information corresponding to the first statement data, contact information, and user identity information. This method utilizes game account information and message data to perform risk assessment and anomaly identification on game accounts, providing users with a safe gaming environment and helping to improve the user's gaming experience.

[0010] Other features and advantages of the present invention are described in a later specification, or partial features and advantages are inferred or uniquely determined from the specification, or can be learned by practicing the above-described techniques of the present invention.

[0011] To further clarify the above-mentioned objectives, features, and advantages of the present invention, preferred embodiments are listed below and described in detail with reference to the accompanying drawings. [Brief explanation of the drawing]

[0012] To more clearly describe specific embodiments of the present invention or technical concepts in the prior art, the following briefly introduces the drawings that may be used in the description of specific embodiments or the prior art. Clearly, the drawings in the following description are some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these without any creative effort. [Figure 1] This is a flowchart of an information identification method according to an embodiment of the present invention. [Figure 2] This is a flowchart for information identification according to an embodiment of the present invention. [Figure 3] This is a schematic diagram of the structure of an information identification device according to an embodiment of the present invention. [Figure 4] This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. [Modes for carrying out the invention]

[0013] To further clarify the object, technical proposal and advantages of the embodiments of the present invention, the technical proposal in the embodiments of the present invention will be clearly and completely described below with reference to the drawings of the embodiments of the present invention. Obviously, the embodiments described are some embodiments of the present invention, not all embodiments. In general, the components in the embodiments of the present invention described and illustrated in the drawings of this specification can be arranged and designed in a variety of different configurations.

[0014] Therefore, the detailed description of embodiments of the present invention provided in the drawings is not intended to limit the scope of the claims, but merely to illustrate specific embodiments of the present invention. Based on the embodiments of the present invention, those skilled in the art will know that all other embodiments obtained without creative work also fall within the scope of the present invention.

[0015] In some popular games, there is a phenomenon where some unscrupulous individuals lure minors to third-party platforms to engage in illegal activities. This phenomenon negatively impacts the games, severely damaging their credibility and leading to user attrition. However, overseas markets have strict demands for the protection of privacy and personal information. Overseas game companies cannot obtain users' identity information, nor do they have the right to arbitrarily investigate users' statements. Furthermore, unscrupulous individuals often lure minors to third-party platforms, making it difficult to leave evidence of illegal activity, and game companies are unable to obtain sufficient leads.

[0016] Based on the above problems, embodiments of the present invention provide an information identification method, apparatus, and electronic device, which can be applied to any game information identification scenario, particularly to information identification scenarios for fraudulent game accounts.

[0017] To facilitate understanding of the embodiments of the present invention, the information identification method disclosed in the embodiments of the present invention will first be described in detail, and as shown in Figure 1, this method includes the following specific steps.

[0018] In step S102, the first statement data corresponding to the first game account in the target game is obtained, abnormal content identification processing is performed on the first statement data, and the identification result is obtained.

[0019] In concrete implementation, the target game may be any game, the first game account may be any game account registered with the target game, and the first statement data corresponding to the first game account includes all statement data sent by the user in the target game through the first game account, and such statement data includes game comments, conversations within a game match, or game messages.

[0020] In actual applications, when a user sends speech data in the target game using a first game account, the speech data is immediately retrieved, abnormal content identification processing is performed on the speech data, and an identification result corresponding to the speech data is obtained. This identification result is used to indicate whether the speech data contains abnormal content or not. Specifically, when identifying abnormal content in speech data, the speech data is compared with keywords in a pre-configured keyword library. If the speech data matches a keyword in the keyword library, the identification result may indicate that the speech data contains abnormal content. Alternatively, the speech data may be matched with a pre-configured regular expression. If both match, the identification result indicates that the speech data contains abnormal content.

[0021] In step S104, the identification result indicates that the first statement data contains abnormal content, the first statement data is stored in a pre-configured database, the account information of the first game account is obtained, and the account information of the first game account is associated with the first statement data and stored in the pre-configured database.

[0022] In a specific implementation, if the identification result indicates that the first speech data does not contain abnormal content, no processing is performed on the first speech data. If the identification result indicates that the first speech data contains abnormal content, it is prepared to further perform manual review and model analysis processing without storing the first speech data in a preset database. It is necessary to store the first speech data in a preset database and determine whether the account information of the first game account corresponding to the first speech data is included in the preset database. If the account information of the first game account is included, it is necessary to explain that the speech data corresponding to the first game account was included in the preset database, directly associate the first speech data with the account information of the first game account, and store them in the preset database. If the account information of the first game account is not included in the preset database, it is necessary to obtain the account information of the first game account from the target game, associate the account information of the first game account with the first speech data, and store them in the preset database.

[0023] Specifically, the above preset database may be a MySQL database, a relational database, a DB database, etc., and may be specifically determined according to development needs. The account information of the first game account includes at least one piece of information such as the number of speeches in the friend channel of the first game account, the number of speeches in the private message channel of strangers, the number of speeches in the team channel, the number of speeches in the room channel, the number of speeches in other channels, the number of games, the number of hits of masked words, the number of adding others to the blacklist, the number of being added to the blacklist, the number of friend applications, the number of friend deletions, the number of deleted friends, the number of moves from the blacklist, the number of received friend applications, the number of times of being punished, the account creation time, and the number of IPs used by the account.

[0024] In step S106, the account information of the first game account in a preset database is input into a pre-trained first model to obtain a first output result. Here, the first output result is used to indicate the probability that the first game account is an abnormal account.

[0025] If the account information of the first game account is stored in a preset database, it is necessary to input the account information of the first game account into a pre-trained first model. The first model can capture the behavior characteristics of the user using the first game account based on the account information of the first game account. Thereby, the first model can predict the probability that the first game account is an abnormal account. Here, the training process of the first model will be described in subsequent embodiments.

[0026] In step S108, when the probability corresponding to the first output result is greater than a preset probability threshold, the first speech data in a preset database is input into a pre-trained second model to obtain a second output result. Here, the second output result includes at least one of the information that the speech is normal, abnormal information corresponding to the first speech data, contact information, and the identity information of the user.

[0027] In concrete implementation, the above-mentioned pre-set probability threshold may be determined according to development needs. For example, the pre-set probability threshold may be set to 80% or 90%. If the probability corresponding to the first output result of the first model is less than or equal to the pre-set probability threshold, the first game account is described as a normal game account, and no processing is performed on the first game account. If the probability corresponding to the first output result is greater than the pre-set probability threshold, the first game account is described as an abnormal account. Consequently, all utterance data corresponding to the first game account in the pre-set database must be input into the second model. The second model performs natural language recognition on the input utterance data to extract text features from the input utterance data, determines whether the input utterance data is an abnormal utterance based on the text features, and if the input utterance data is an abnormal utterance, obtains one or more of the following from the input utterance data: abnormal information, contact information, and user identity information.

[0028] The user's identity information described above can be used to indicate whether the user using the first game account is a minor, or whether the utterance data contains information related to a minor. In practical applications, the abnormal content in the utterance data is determined based on the second output result, and processing such as masking is applied to the abnormal content to provide a safer language and game environment for minors, thereby ensuring their safety in games.

[0029] An embodiment of the present invention provides an information identification method that utilizes account information and message data of a game account to perform risk assessment and anomaly identification on the game account, thereby providing a safe gaming environment for the user and helping to improve the user's gaming experience.

[0030] The following examples illustrate a method for identifying anomalies in spoken data.

[0031] Specifically, the process for performing abnormal content identification processing on the first statement data mentioned above and obtaining the identification result includes at least one of the following:

[0032] Firstly, the first statement data is matched with a pre-configured regular expression library, the first matching result is obtained, and the first matching result is determined to be the identification result corresponding to the first statement data. Here, if the first statement data matches the regular expression in the regular expression library, the first matching result indicates that the first statement data contains abnormal content.

[0033] In practical implementation, the regular expressions in the regular expression library are pre-configured according to development needs, and the keywords that each regular expression can match can be varied in many ways. For example, the regular expression " / live.*? / " can match variations of keywords such as "where do you live?" and "where do you live?", thus broadening the coverage when performing keyword matching using regular expressions.

[0034] Secondly, the first statement data is matched with a pre-configured keyword library to obtain a second matching result, and the second matching result is determined as the identification result corresponding to the first statement data. Here, if the first statement data hits a keyword in the keyword library, the second matching result indicates that the first statement data contains abnormal content.

[0035] The keyword library mentioned above contains many keywords, and these keywords may be determined according to development needs. For example, the keywords may be "sex" or "where to go." In a concrete implementation, if the first utterance data contains keywords from the keyword library, or if the first utterance data contains a predetermined number of keywords, it is determined that the first utterance data is a match for a keyword in the keyword library.

[0036] In one specific embodiment, in order to cover more scenes and speed up matching without affecting the user experience, the present invention can perform abnormal matching simultaneously using a regular expression library and a keyword library. Figure 2 is a flowchart of information identification according to an embodiment of the present invention, and the regular keyword library matching process in Figure 2 includes the simultaneous passage of a regular matching module and a keyword search module in the SDK (Software Development Kit) when a user issues utterance data in the target game. Here, the regular matching module matches the utterance data using a regular expression library, and the keyword search module matches the utterance data using a keyword library. Here, the regular matching module can increase keyword coverage, and the keyword search module is responsible for quickly searching for abnormal content. As a result, the method can quickly and efficiently block known risk signals (which correspond to utterance data containing abnormal content). Finally, the utterance data suspected of being abnormal is stored in a pre-configured database in preparation for further manual review and model analysis. Here, the utterance data suspected of being abnormal is utterance data matched to a regular expression in the regular expression library, or utterance data that hits a keyword in the keyword library.

[0037] Specifically, the keyword search algorithm is low in complexity (O(1)), can quickly match known anomalies, and does not require complex pattern matching by directly searching a table, resulting in extremely fast search speeds and effectively covering anomaly scenes with clear intent. For example, the keyword "sex" can be matched to a specific anomaly scene. Since it is difficult to cover all variants and implicit anomaly terms, relying solely on a keyword library is insufficient. Therefore, this invention introduces a regular expression library to compensate for the shortcomings of the keyword library.

[0038] In selectable embodiments, the method provided by the present invention is used on a server, but the SDK's regular matching module can be placed on the game client and the keyword module on the server to perform searches without the user being able to detect the search matching. Therefore, the regular matching module placed on the game client corresponding to the target game matches the first utterance data with a pre-configured regular expression library, obtains the first matching result, and sends the first matching result to the server. In addition, the regular matching module can obtain more utterance data by intercepting messages that the user clicked and sent in the target game but were not sent.

[0039] The above method can quickly and accurately identify whether or not the message data contains abnormal content. Based on this, the game account corresponding to the message data containing abnormal content can then be further determined to determine whether or not the game account is an abnormal account.

[0040] The following example is used to illustrate the training method for the first model.

[0041] Specifically, the first model described above is obtained through training in steps 10-14 below.

[0042] In step 10, a first training set is obtained, which includes multiple first training samples, each of which includes a first pre-configured account information and a first label corresponding to the first pre-configured account information, the first label being used to indicate that the first pre-configured account information is abnormal information.

[0043] In concrete implementation, the first training sample in the first training set described above is derived from accused and verified user data, that is, the user data is account information corresponding to a game account accused and verified to have abnormal behavior, the user data corresponds to the first pre-configured account information described above, and since the user data is account information of a game account in which abnormal behavior exists, the first label corresponding to the first pre-configured account information is used to indicate that the first pre-configured account information is abnormal information.

[0044] In step 11, the initial model is trained based on the first training set, and an intermediate model is obtained.

[0045] In concrete implementation, first, a target first training sample can be selected from the first training set. First, the first pre-configured account information from the target first training sample is input into the initial model, and the model output result is obtained. Then, the model parameters of the initial model are adjusted based on the model output result and the first label from the target first training sample, and the initial model after parameter adjustment is obtained. The step of selecting a target first training sample from the first training set is continued until the initial model after parameter adjustment converges, and the converged initial model is determined to be the intermediate model. Here, each first training sample in the first training set can participate in model training as a target first training sample once.

[0046] In step 12, a second training set is obtained, which includes multiple second training samples, and each second training sample includes second pre-configured account information.

[0047] The second training sample included in the second training set described above is a second pre-configured account information that does not have a label set, and this second pre-configured account information may be the account information of any game account obtained from the target game. In other words, since the second pre-configured account information is account information that has not been verified as to whether or not it is an abnormal account, it is not possible to automatically set a label on the second pre-configured account information.

[0048] In step 13, the second training sample from the second training set is input into the intermediate model, the model output is obtained, and a second label is set for the second training sample based on the model output. The second label is used to indicate the probability that the second pre-set account information in the second training sample is anomaly information.

[0049] In practical implementation, because the amount of user data for which accusations have been made and fact-checked is small, pseudo-labeling techniques can be used to fully utilize the acquired behavioral feature data to set labels for a second training set, enabling the construction of a high-quality model that can accurately capture the essential disciplines of user behavioral characteristics and have good generalization ability for unseen data.

[0050] In practical applications, the second training sample from the second training set can be input into the intermediate model, the model output can be obtained, and then either the model output can be used as the second label for the second training sample, or the product of the model output and a pre-set weight can be used as the second label for the second training sample.

[0051] In step 14, an intermediate model is trained based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set, and the first model is obtained.

[0052] In concrete implementation, an intermediate model can be trained based on all first training samples in the first training set, some or all second training samples in the second training set, and second labels corresponding to the second training samples, thereby obtaining the first model.

[0053] In an optional embodiment, a second training sample is selected from the second training set in which the second label is greater than a preset value. The selected second training sample and the second label corresponding to the selected second training sample are used as the third training sample. An intermediate model can be trained based on the third training set and the first training set to obtain the first model. The preset value may be determined according to development needs; for example, the preset value may be 80% or 90%.

[0054] In another optional embodiment, the second training samples may be sorted in descending order of their probability of corresponding to the second label, and the top specified number of second training samples and the second labels corresponding to those second training samples may be used as the third training sample.

[0055] In one specific embodiment, the initial model may be a machine learning model or a neural network model, for example, an XGBoost model, where the XGBoost model has parallelization features and can generate excellent models quickly and efficiently.

[0056] In another optional embodiment, a single initial model may overfit the training data, thus reducing its ability to adapt to new, unseen data. Therefore, the initial model in the present invention may include multiple submodels, thereby training each submodel based on a first training set to obtain multiple intermediate submodels, further inputting a second training sample from a second training set into each intermediate submodel, obtaining the output result of each intermediate submodel, determining a second label for the second training sample based on the output result of each intermediate submodel, then training each intermediate submodel based on the second label corresponding to the first training set, the second training set, and the second training sample in the second training set, obtaining multiple trained intermediate submodels, and determining the multiple trained intermediate submodels as the first model. Here, the training method for each submodel in the initial model is the same, and different submodels are trained individually.

[0057] In concrete implementation, the specific process for determining the second label of the second training sample based on the output results of each of the above intermediate submodels may include any of the following:

[0058] First, the average of the output results of each intermediate submodel is determined as the second label for the second training sample. For example, if there are three intermediate submodels, the second label for the second training sample is the sum of the output results of the three intermediate submodels divided by three.

[0059] Secondly, a second label for the second training sample is determined based on the output results of each intermediate submodel and the output weights corresponding to each intermediate submodel, where the output weights corresponding to each intermediate submodel are determined based on a second pre-configured account information for the second training sample.

[0060] In concrete implementation, the information content corresponding to the second pre-configured account information differs, the output weights corresponding to each intermediate submodel also differ, and which information content in the second pre-configured account information affects the output weight corresponding to the intermediate submodel is predetermined. For example, if there are three intermediate submodels, and the number of posts in the friend channel in the second pre-configured account information is large, the output weight corresponding to the first intermediate submodel will be high, and the output weights corresponding to the other two intermediate submodels will be low. If the number of deleted friends is large, the output weight corresponding to the second intermediate submodel will be high.

[0061] In one specific example, the initial model shown in Figure 2 includes three boosting models: XGBoost, CatBoost, and LightGBM. These three submodels are trained using the first training set to obtain intermediate submodels corresponding to these three submodels. These three intermediate submodels are then applied to a large number of unlabeled data samples (corresponding to the second training set), and prediction and pseudo-labeling are performed on them. Next, the portion of these pseudo-labeled data (corresponding to the second training sample and the second label corresponding to the second training sample) with the highest model prediction probability is selected and considered as additional labeled samples (corresponding to the third training set). These are added to the first training set, and based on the third and first training sets, these three intermediate submodels are trained to obtain three trained intermediate submodels.

[0062] Furthermore, to further improve the robustness of the models, a large amount of sample data augmented with pseudo-labeling techniques can be trained in parallel using 5-fold cross-validation (5-CV) with three different boosting tree models: XGBoost, CatBoost, and LightGBM. Finally, the output probabilities from the 5-fold validation of the three models are stacked, meaning the output probabilities of the three models are used as new features, and these probabilistic features are integrated to obtain the final predicted probability (corresponding to the second label mentioned above).

[0063] Based on the above description, the pre-trained first model includes a plurality of trained intermediate submodels, and the specific process of inputting the account information of the first game account in the pre-configured database into the pre-trained first model and obtaining a first output result may include inputting the account information of the first game account in the pre-configured database into each of the trained intermediate submodels and obtaining an output result corresponding to each trained intermediate submodel, determining the average value of the output results corresponding to each trained intermediate submodel as the first output result, or determining the first output result based on the output result corresponding to each trained intermediate submodel and the output weights corresponding to each trained intermediate submodel, wherein the output weights corresponding to the trained intermediate submodels are determined based on the information contained in the account information of the first game account.

[0064] The self-training method described above successfully expands the size and diversity of the training set, bringing the data distribution the model saw during training closer to real-world conditions, thereby effectively improving the model's ability to generalize and adapt to unfamiliar data.

[0065] The following example is used to illustrate the training method for the second model.

[0066] Specifically, the second model described above is obtained by training as follows, acquiring a fourth training set, the fourth training set containing multiple fourth training samples, the fourth training samples containing pre-configured utterance data and a fourth label corresponding to the pre-configured utterance data, the fourth label being used to indicate whether the pre-configured utterance data is an abnormal utterance, and if the pre-configured utterance is an abnormal utterance, it contains at least one piece of information from the abnormal information, contact information, and user identity information corresponding to the pre-configured utterance data, the second model is trained based on the fourth training set and a pre-configured dynamic learning rate scheduling policy, and the trained second model is acquired.

[0067] In concrete implementation, the pre-configured speech data in the fourth training sample of the fourth training set is game speech data corresponding to user data that has been reported to have detected an anomaly, and the number of training samples included in the fourth training set may be determined according to development needs, for example, the number of training samples may be 1000 or 800. After obtaining the pre-configured speech data in the fourth training sample, it is necessary to manually label the pre-configured speech data, and a fourth label corresponding to the pre-configured speech data is obtained, where the fourth label is used to indicate whether the pre-configured speech data is an anomaly, and if the pre-configured speech is an anomaly, it includes at least one of the anomaly information, contact information and user identity information corresponding to the pre-configured speech data.

[0068] In one specific implementation, when manually labeling, four types of labels can be applied to pre-configured speech data: one that specifies the user's identity, one that relates to an abnormal scene, one that includes contact information and normal content (corresponding to the above speech not being abnormal). Contextual information is also tightly integrated into the labeling process, as it is often necessary to combine context to accurately understand the meaning of certain words or phrases in the pre-configured speech data. For example, for the Japanese phrase "yaru?" (do? or play?), the fourth label in the labeling result will differ depending on the surrounding context.

[0069] In concrete implementation, to improve the convergence speed and fitting ability during the model training process, the second model may be optimized based on LORA. First, instead of using a unified global learning rate, the learning rate parameters of each layer of the second model's network can be individually set based on LORA technology. This method allows different layers to adopt different learning rates according to their own characteristics, making network parameter updates more precise and purposeful. Next, the present invention introduces a dynamic learning rate scheduling policy to monitor the changes in the loss function during the training process in real time and automatically adjust the learning rate of each layer. If the loss function tends to slow down, the learning rate is appropriately increased to speed up convergence, and if the loss function fluctuates sharply, the learning rate is lowered to avoid oscillations and divergences. This dynamic adjustment mechanism contributes to capturing the optimal convergence point and significantly improves the convergence efficiency of the second model within a limited number of training steps.

[0070] In concrete implementation, the second model described above may employ a large-scale language model (which may also be called an LLM model), or other neural network models, and may be determined specifically according to development needs. Herein, the present invention may employ the Llama2 large-scale model in the LLM model, which is a general-purpose large-scale model capable of identifying 150 languages ​​such as Japanese, Chinese, and English. During the training process, three pieces of data need to be input: Prompt is a prompt word with the content "You are a valid language assistant and can identify abnormal speech. Please make a judgment on the following data: Label is normal, User identity information, Contact information, Abnormal information, User identity information + Abnormal information, User identity information + Contact, User identity information + Abnormal information + Contact, Abnormal information + Contact"; text is "Hello, I'm xxx, would you like to be friends with me? I'm a 16-year-old student, where are you from? Hey, would you like to hang out?"; and output is "User identity information + Contact," where the User identity information may be the identity of a user whose age is below a pre-set age threshold, and this pre-set age threshold may be determined according to development needs.

[0071] In the selectable embodiment, the pre-set speech data included in the fourth training sample is: Sliding window The original utterance data is obtained by dividing it, and when the original utterance data is divided, there is an overlap between the earlier and later parts of the divided utterance data.

[0072] In its specific implementation, the present invention performs a rational segmentation process on labeled training samples. Experiments have shown that if text segments (corresponding to pre-set utterance data) are too long or too short, it affects the accurate identification and understanding of text content by the large-scale language model (corresponding to the second model described above). Therefore, the present invention provides a rational segmentation process for labeled training samples. Sliding window The following method is used to divide the pre-set speech data: Sliding window By appropriately adjusting the length, sufficient text length is maintained while expanding the scope of the context as much as possible, allowing the second model to better capture contextual information and improving its learning ability for text content. Next, to avoid misunderstandings due to loss of context, an overlap mechanism is introduced, that is, two adjacent elements Sliding window By maintaining a certain degree of overlap between them, the divided preceding and succeeding speech data will have overlapping portions, ensuring that important contextual information can exist simultaneously in adjacent windows.

[0073] In selectable embodiments, the second model can be compressed and combined to enable smooth deployment of the second model to the server. Specifically, the trained second model can be fine-tuned, the fine-tuned second model can be obtained, the fine-tuned second model can be compressed and combined to obtain the final second model, and the final second model can be used to predict speech data.

[0074] In practical implementation, the conventional LORA (Low-Rank Adaptation) fine-tuning process updates only a small amount of offset and attention parameters in the second model, leaving most of the parameters of the pre-trained second model unchanged. While this approach can accelerate training and save memory, it requires loading the full model parameters of the pre-trained second model during the inference phase, resulting in significant extra memory and computational overhead. Therefore, QLORA (Quantized Low-Rank Adaptation) technology can be used to compress and combine the fine-tuned parameters of LORA to generate dedicated "inference weights," and the specific process is as follows.

[0075] 1. Derive and calculate an equivalent full-parameter matrix based on a small offset after updating LORA and the attention parameter.

[0076] 2. These full-parameter matrices are compressed, reducing the parameter dimension and memory occupancy through techniques such as matrix decomposition and quantization.

[0077] 3. The compressed parameters are combined with the original parameters of a second, pre-trained model to generate a new inference weight tensor.

[0078] According to the above method, risk assessment and anomaly identification are performed on users by comprehensively utilizing user account information and speech data. Intelligent modeling and big data analysis enable the efficient detection of users suspected of anomaly. To address the specific needs of protecting minors, this method successfully trains a highly accurate multi-label language type classification model by combining a small labeling corpus with large-scale language modeling techniques, significantly reducing human costs. This model can accurately identify inappropriate content across various language types, providing a safer language environment for minors.

[0079] Corresponding to the embodiments of the above method, an embodiment of the present invention further provides an information identification device, as shown in Figure 3, the device comprises a data acquisition module 30, an information storage module 31, a first identification module 32, and a second identification module 33.

[0080] The data acquisition module 30 is configured to acquire first statement data corresponding to a first game account in the target game, perform abnormal content identification processing on the first statement data, and acquire the identification result. The information storage module 31 is configured to store the first statement data in a pre-configured database if the identification result indicates that the first statement data contains abnormal content, retrieve the account information of the first game account, and store the account information of the first game account in a pre-configured database in association with the first statement data. The first identification module 32 is configured to input the account information of the first game account from a pre-configured database into a pre-trained first model and obtain a first output result, which is used to indicate the probability that the first game account is an abnormal account. The second identification module 33 is configured to input the first utterance data from a pre-configured database into a pre-trained second model and obtain a second output result if the probability corresponding to the first output result is greater than a pre-configured probability threshold, the second output result includes at least one of the following: that the utterance is normal, anomaly information corresponding to the first utterance data, contact information, and user identity information.

[0081] According to the information identification device described above, this method utilizes the account information and message data of game accounts to perform risk assessment and anomaly identification on game accounts, thereby providing users with a safe gaming environment and contributing to improving the user's gaming experience.

[0082] Specifically, the data acquisition module 30 matches the first utterance data with a pre-configured regular expression library, obtains a first matching result, and determines the first matching result as the identification result corresponding to the first utterance data. If the first utterance data matches a regular expression in the regular expression library, the first matching result indicates that the first utterance data contains abnormal content. And / or, the module matches the first utterance data with a pre-configured keyword library, obtains a second matching result, and determines the second matching result as the recognition result corresponding to the first utterance data. If the first utterance data hits a keyword in the keyword library, the second matching result indicates that the first utterance data contains abnormal content.

[0083] In concrete implementation, the above method is applied to the server, and the data acquisition module 30 of the device is further configured to match the first spoken data with a pre-configured regular expression library using a regular matching module located in the game client corresponding to the target game, obtain the first matching result, and send the first matching result to the server.

[0084] Furthermore, the device further includes a first model training module for acquiring a first training set, the first training set including a plurality of first training samples, the first training samples including first pre-configured account information and a first label corresponding to the first pre-configured account information, the first label being used to indicate that the first pre-configured account information is abnormal information, training an initial model based on the first training set, acquiring an intermediate model, acquiring a second training set, the second training set including a plurality of second training samples The second training sample includes a second set of pre-configured account information. The second training sample in the second training set is input into the intermediate model, the model output is obtained, and a second label is set for the second training sample based on the model output. The second label is used to indicate the probability that the second set of pre-configured account information in the second training sample is anomaly information. The intermediate model is trained based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set, and the first model is obtained.

[0085] In concrete implementation, the first model training module is further configured to select a second training sample from the second training set in which the second label is greater than a preset value, to make the selected second training sample and the second label corresponding to the selected second training sample a third training sample, and to train an intermediate model based on the third training set and the first training set to obtain the first model.

[0086] In selectable embodiments, the initial model includes a plurality of submodels, the first model training module is configured to further train each submodel based on a first training set to obtain a plurality of intermediate submodels, input a second training sample from a second training set into each intermediate submodel, obtain the output result of each intermediate submodel, determine a second label for the second training sample based on the output result of each intermediate submodel, train each intermediate submodel based on the first training set, the second training set and the second label corresponding to the second training sample in the second training set, obtain a plurality of trained intermediate submodels, and determine the plurality of trained intermediate submodels as the first model.

[0087] Furthermore, determining a second label for a second training sample based on the output results of each intermediate submodel includes either determining the average value of the output results of each intermediate submodel as the second label for the second training sample, or determining a second label for a second training sample based on the output results of each intermediate submodel and the output weights corresponding to each intermediate submodel, wherein the output weights corresponding to each intermediate submodel are determined based on a second set of pre-configured account information in the second training sample.

[0088] In concrete implementation, the first model includes a plurality of trained intermediate submodels, and the first identification module 32 inputs the account information of the first game account in a pre-configured database to each of the trained intermediate submodels, obtains the output result corresponding to each trained intermediate submodel, determines the average value of the output results corresponding to each trained intermediate submodel as the first output result, or determines the first output result based on the output result corresponding to each trained intermediate submodel and the output weights corresponding to each trained intermediate submodel, where the output weights corresponding to the trained intermediate submodels are configured to be determined based on the information contained in the account information of the first game account.

[0089] Furthermore, the device further includes a second model training module for acquiring a fourth training set, the fourth training set including a plurality of fourth training samples, the fourth training samples including pre-configured utterance data and a fourth label corresponding to the pre-configured utterance data, the fourth label being used to indicate whether the pre-configured utterance data is an abnormal utterance, and if the pre-configured utterance is an abnormal utterance, including at least one of the following: abnormal information, contact information, and user identity information corresponding to the pre-configured utterance data, the second model is trained based on the fourth training set and a pre-configured dynamic learning rate scheduling policy, and the trained second model is acquired.

[0090] In the selectable embodiment, the pre-set speech data included in the fourth training sample is: Sliding window The original utterance data is obtained by dividing it, and when the original utterance data is divided, there is an overlap between the earlier and later parts of the divided utterance data.

[0091] Furthermore, the above device further includes a model compression module, which is configured to perform fine-tuning on the trained second model, obtain the fine-tuned second model, perform compression and combination processing on the fine-tuned second model, obtain the final second model, and use the final second model to predict speech data.

[0092] The information identification device according to the embodiment of this disclosure has the same implementation principle and generated technical effects as the embodiment of the method described above. A brief explanation is provided, and for parts that do not refer to the embodiment of the device, the corresponding content in the embodiment of the method described above can be referenced.

[0093] Embodiments of the present invention further provide an electronic device, as shown in Figure 4, which includes a processor and a memory, the memory storing device-executable instructions that can be executed by the processor, and the processor executing the device-executable instructions to realize the information identification method.

[0094] Specifically, the above information identification method includes: obtaining first statement data corresponding to a first game account in the target game; performing abnormal content identification processing on the first statement data and obtaining an identification result; if the identification result indicates that the first statement data contains abnormal content, storing the first statement data in a pre-configured database; obtaining account information for the first game account; associating the account information for the first game account with the first statement data and storing it in a pre-configured database; inputting the account information for the first game account in the pre-configured database into a pre-trained first model and obtaining a first output result, wherein the first output result is used to indicate the probability that the first game account is an abnormal account; and if the probability corresponding to the first output result is greater than a pre-configured probability threshold, inputting the first statement data in the pre-configured database into a pre-trained second model and obtaining a second output result, wherein the second output result includes at least one of the following: that the statement is normal, abnormal information corresponding to the first statement data, contact information, and user identity information.

[0095] The above information identification method utilizes the account information and message data of game accounts to perform risk assessment and anomaly identification on game accounts, thereby providing users with a safe gaming environment and contributing to improving the user's gaming experience.

[0096] In an optional embodiment, performing abnormal content identification processing on the first statement data and obtaining an identification result includes at least one of the following: matching the first statement data with a pre-configured regular expression library, obtaining a first matching result, and determining the first matching result as the identification result corresponding to the first statement data, wherein if the first statement data matches a regular expression in the regular expression library, the first matching result indicates that the first statement data contains abnormal content; and matching the first statement data with a pre-configured keyword library, obtaining a second matching result, and determining the second matching result as the recognition result corresponding to the first statement data, wherein if the first statement data hits a keyword in the keyword library, the second matching result indicates that the first statement data contains abnormal content.

[0097] In an optional embodiment, the method is applied to a server, and further includes matching first utterance data with a pre-configured regular expression library using a regular matching module located in a game client corresponding to the target game, obtaining a first matching result, and sending the first matching result to the server.

[0098] In an optional embodiment, the first model is obtained by training as follows: a first training set is obtained, the first training set includes a plurality of first training samples, the first training samples include first pre-configured account information and a first label corresponding to the first pre-configured account information, the first label is used to indicate that the first pre-configured account information is abnormal information, an initial model is trained based on the first training set, an intermediate model is obtained, a second training set is obtained, the second training set includes a plurality of second training The first training set includes a sample, and the second training sample includes a second set of pre-configured account information. The second training sample in the second training set is input into the intermediate model, and the model output is obtained. Based on the model output, a second label is set for the second training sample. The second label is used to indicate the probability that the second set of pre-configured account information in the second training sample is anomaly information. The intermediate model is trained based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set, and the first model is obtained.

[0099] In an optional embodiment, training an intermediate model and obtaining a first model based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set includes selecting a second training sample from the second training set in which the second label is greater than a preset value, making the selected second training sample and the second label corresponding to the selected second training sample a third training sample, and training an intermediate model and obtaining a first model based on the third training set and the first training set.

[0100] In selectable embodiments, the initial model includes a plurality of submodels, and training the initial model based on the first training set and obtaining an intermediate model includes training each submodel based on the first training set and obtaining a plurality of intermediate submodels, inputting a second training sample from the second training set into the intermediate models, obtaining the model output results, and setting a second label for the second training sample based on the model output results, inputting a second training sample from the second training set into each intermediate submodel, obtaining the output results for each intermediate submodel The process of obtaining the first model includes determining a second label for a second training sample based on Dell's output, training an intermediate model based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set, training each intermediate submodel based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set, obtaining a set of trained intermediate submodels, and determining the set of trained intermediate submodels as the first model.

[0101] In selectable embodiments, determining a second label for a second training sample based on the output results of each intermediate submodel includes determining the average value of the output results of each intermediate submodel as the second label for the second training sample, or determining a second label for a second training sample based on the output results of each intermediate submodel and the output weights corresponding to each intermediate submodel, wherein the output weights corresponding to the intermediate submodels are determined based on second pre-configured account information in the second training sample.

[0102] In an optional embodiment, the first model includes a plurality of trained intermediate submodels, and inputting account information of a first game account in a pre-configured database into a pre-trained first model to obtain a first output result includes inputting account information of a first game account in a pre-configured database into each of the trained intermediate submodels and obtaining an output result corresponding to each trained intermediate submodel, determining the average value of the output results corresponding to each trained intermediate submodel as the first output result, or determining the first output result based on the output result corresponding to each trained intermediate submodel and the output weights corresponding to each trained intermediate submodel, wherein the output weights corresponding to the trained intermediate submodels are determined based on the information contained in the account information of the first game account.

[0103] In an optional embodiment, the second model is obtained by training as follows to acquire a fourth training set, the fourth training set containing a plurality of fourth training samples, the fourth training samples containing pre-configured utterance data and a fourth label corresponding to the pre-configured utterance data, the fourth label being used to indicate whether the pre-configured utterance data is an abnormal utterance, and if the pre-configured utterance is an abnormal utterance, it contains at least one of the following: abnormal information corresponding to the pre-configured utterance data, contact information, and user identity information, the second model is trained based on the fourth training set and a pre-configured dynamic learning rate scheduling policy, and the trained second model is acquired.

[0104] In the selectable embodiment, the pre-set speech data included in the fourth training sample is: Sliding window The original utterance data is obtained by dividing it, and when the original utterance data is divided, there is an overlap between the earlier and later parts of the divided utterance data.

[0105] In an optional embodiment, the method further includes fine-tuning the trained second model to obtain a fine-tuned second model, compressing and combining the fine-tuned second model to obtain a final second model, and using the final second model to predict speech data.

[0106] Furthermore, the electronic device shown in Figure 4 further includes a bus 102 and a communication interface 103, and the processor 101, communication interface 103, and memory 100 are connected via the bus 102.

[0107] Here, memory 100 may include high-speed random access memory (RAM), and may further include non-volatile memory, such as at least one magnetic disk memory. At least one communication interface 103 (which may be wired or wireless) enables communication between the system network element and at least one other network element, allowing the use of the internet, wide area network, local network, metro network, etc. Bus 102 may be an ISA bus, PCI bus, EISA bus, etc. The bus can be divided into an address bus, data bus, control bus, etc. For ease of representation, only one double-headed arrow is shown in Figure 4, but this does not mean that only one bus or one type of bus is shown.

[0108] The processor 101 may be an integrated circuit chip having signal processing capabilities. In the implementation process, each step of the above method can be completed by hardware integrated logic circuits or software-based instructions in the processor 101. The processor 101 may be a general-purpose processor such as a Central Processing Unit (CPU), a Network Processor (NP), a Digital Signal Processing (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, or discrete hardware components. Each method, step, and logic block diagram disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present invention may be performed directly by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be installed in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, or registers. The storage medium is installed in memory 100, and the processor 101 reads the information in memory 100 and, in combination with its hardware, completes the steps of the method of the embodiment.

[0109] Embodiments of the present invention further provide a computer-readable storage medium in which computer-executable instructions are stored, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the information identification method described above. A detailed explanation of the implementation is provided by referring to the embodiments of the method and is omitted here.

[0110] Specifically, the above information identification method includes: obtaining first statement data corresponding to a first game account in the target game; performing abnormal content identification processing on the first statement data and obtaining an identification result; if the identification result indicates that the first statement data contains abnormal content, storing the first statement data in a pre-configured database; obtaining account information for the first game account; associating the account information for the first game account with the first statement data and storing it in a pre-configured database; inputting the account information for the first game account in the pre-configured database into a pre-trained first model and obtaining a first output result, wherein the first output result is used to indicate the probability that the first game account is an abnormal account; and if the probability corresponding to the first output result is greater than a pre-configured probability threshold, inputting the first statement data in the pre-configured database into a pre-trained second model and obtaining a second output result, wherein the second output result includes at least one of the following: that the statement is normal, abnormal information corresponding to the first statement data, contact information, and user identity information.

[0111] The above information identification method utilizes the account information and message data of game accounts to perform risk assessment and anomaly identification on game accounts, thereby providing users with a safe gaming environment and contributing to improving the user's gaming experience.

[0112] In an optional embodiment, performing abnormal content identification processing on the first statement data and obtaining an identification result includes at least one of the following: matching the first statement data with a pre-configured regular expression library, obtaining a first matching result, and determining the first matching result as the identification result corresponding to the first statement data, wherein if the first statement data matches a regular expression in the regular expression library, the first matching result indicates that the first statement data contains abnormal content; and matching the first statement data with a pre-configured keyword library, obtaining a second matching result, and determining the second matching result as the recognition result corresponding to the first statement data, wherein if the first statement data hits a keyword in the keyword library, the second matching result indicates that the first statement data contains abnormal content.

[0113] In an optional embodiment, the method is applied to a server, and further includes matching first utterance data with a pre-configured regular expression library using a regular matching module located in a game client corresponding to the target game, obtaining a first matching result, and sending the first matching result to the server.

[0114] In an optional embodiment, the first model is obtained by training as follows: a first training set is obtained, the first training set includes a plurality of first training samples, the first training samples include first pre-configured account information and a first label corresponding to the first pre-configured account information, the first label is used to indicate that the first pre-configured account information is abnormal information, an initial model is trained based on the first training set, an intermediate model is obtained, a second training set is obtained, the second training set includes a plurality of second training The first training set includes a sample, and the second training sample includes a second set of pre-configured account information. The second training sample in the second training set is input into the intermediate model, and the model output is obtained. Based on the model output, a second label is set for the second training sample. The second label is used to indicate the probability that the second set of pre-configured account information in the second training sample is anomaly information. The intermediate model is trained based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set, and the first model is obtained.

[0115] In an optional embodiment, training an intermediate model and obtaining a first model based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set includes selecting a second training sample from the second training set in which the second label is greater than a preset value, making the selected second training sample and the second label corresponding to the selected second training sample a third training sample, and training an intermediate model and obtaining a first model based on the third training set and the first training set.

[0116] In selectable embodiments, the initial model includes a plurality of submodels, and training the initial model based on the first training set and obtaining an intermediate model includes training each submodel based on the first training set and obtaining a plurality of intermediate submodels, inputting a second training sample from the second training set into the intermediate models, obtaining the model output results, and setting a second label for the second training sample based on the model output results, inputting a second training sample from the second training set into each intermediate submodel, obtaining the output results for each intermediate submodel The process of obtaining the first model includes determining a second label for a second training sample based on Dell's output, training an intermediate model based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set, training each intermediate submodel based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set, obtaining a set of trained intermediate submodels, and determining the set of trained intermediate submodels as the first model.

[0117] In selectable embodiments, determining a second label for a second training sample based on the output results of each intermediate submodel includes determining the average value of the output results of each intermediate submodel as the second label for the second training sample, or determining a second label for a second training sample based on the output results of each intermediate submodel and the output weights corresponding to each intermediate submodel, wherein the output weights corresponding to the intermediate submodels are determined based on second pre-configured account information in the second training sample.

[0118] In an optional embodiment, the first model includes a plurality of trained intermediate submodels, and inputting the account information of the first game account in the pre-configured database into the pre-trained first model and obtaining a first output result includes inputting the account information of the first game account in the pre-configured database into each of the trained intermediate submodels and obtaining an output result corresponding to each trained intermediate submodel, determining the average value of the output results corresponding to each trained intermediate submodel as the first output result, or determining the first output result based on the output result corresponding to each trained intermediate submodel and the output weights corresponding to each trained intermediate submodel, wherein the output weights corresponding to the trained intermediate submodels are determined based on the information contained in the account information of the first game account.

[0119] In an optional embodiment, the second model is obtained by training as follows to acquire a fourth training set, the fourth training set containing a plurality of fourth training samples, the fourth training samples containing pre-configured utterance data and a fourth label corresponding to the pre-configured utterance data, the fourth label being used to indicate whether the pre-configured utterance data is an abnormal utterance, and if the pre-configured utterance is an abnormal utterance, it contains at least one of the following: abnormal information corresponding to the pre-configured utterance data, contact information, and user identity information, the second model is trained based on the fourth training set and a pre-configured dynamic learning rate scheduling policy, and the trained second model is acquired.

[0120] In the selectable embodiment, the pre-set speech data included in the fourth training sample is: Sliding window The original utterance data is obtained by dividing it, and when the original utterance data is divided, there is an overlap between the earlier and later parts of the divided utterance data.

[0121] In an optional embodiment, the method further includes fine-tuning the trained second model to obtain a fine-tuned second model, compressing and combining the fine-tuned second model to obtain a final second model, and using the final second model to predict speech data.

[0122] The aforementioned functions may be implemented as software function units and, when sold or used as standalone products, may be stored on a computer-readable storage medium. Based on this understanding, it is understood that any part of the proposed technical proposal of the present invention that is essentially or contributes to the prior art may be embodied in the form of a software product stored on a storage medium containing a number of instructions that enable a computer device (which may be a personal computer, terminal device, or network device, etc.) to perform all or part of the methods described in various embodiments of the present invention. Examples of such storage mediums include various media capable of storing program code, such as USB flash drives, removable hard disks, read-only memory (ROM), random access memory (RAM), disks, or CD-ROMs.

[0123] In the description of this invention, directions or positional relationships indicated by terms such as "center," "up," "down," "left," "right," "vertical," "horizontal," "inside," and "outside" are based on the directions or positional relationships shown in the drawings and are merely for the purpose of facilitating and simplifying the description of this invention. They do not indicate or imply that the mentioned devices or components have a specific orientation or must be configured and operated in a specific orientation, and therefore cannot be understood as limiting this invention. Furthermore, the terms "first," "second," and "third" are merely for the purpose of describing the objective and cannot be understood as indicating or implying relative importance.

[0124] Finally, it should be noted that the above embodiments are merely specific embodiments of the present invention and are intended to illustrate the technical proposal of the present invention, but are not limited thereto. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art can still modify or easily alter the technical proposal described in the above embodiments within the technical scope disclosed herein, or make equivalent substitutions to some of the technical features thereof. Such modifications, alterations, or substitutions should be included within the scope of protection of the present invention without departing the essence of the corresponding technical proposal from the spirit and scope of the technical proposal of the embodiments of the present invention. Therefore, the scope of protection of the present invention is equivalent to the scope of protection of the claims.

Claims

1. The process involves obtaining first statement data corresponding to a first game account in the target game, performing abnormal content identification processing on the first statement data, and obtaining the identification result. If the identification result indicates that the first statement data contains abnormal content, the first statement data is stored in a pre-configured database, the account information of the first game account is obtained, and the account information of the first game account is stored in the pre-configured database in association with the first statement data. The method involves inputting the account information of the first game account in the pre-configured database into a pre-trained first model and obtaining a first output result, wherein the first output result is used to indicate the probability that the first game account is an abnormal account. If the probability corresponding to the first output result is greater than a predetermined probability threshold, the first utterance data in the predetermined database is input into a pre-trained second model, and a second output result is obtained, wherein the second output result includes at least one of the following: information indicating that the first utterance data is normal, abnormal information corresponding to the first utterance data, contact information, and user identity information. A method for identifying information characterized by the following features.

2. Performing abnormal content identification processing on the first statement data and obtaining the identification result is: The first statement data is matched with a pre-configured regular expression library, a first matching result is obtained, and the first matching result is determined as an identification result corresponding to the first statement data, wherein if the first statement data is matched with a regular expression in the regular expression library, the first matching result indicates that the first statement data contains abnormal content. The method involves matching the first statement data with a pre-configured keyword library, obtaining a second matching result, and determining the second matching result as the recognition result corresponding to the first statement data, wherein if the first statement data hits a keyword in the keyword library, the second matching result indicates that the first statement data contains abnormal content, and this method includes at least one of these steps. The information identification method according to feature 1.

3. An information identification method applied to a server, The process further includes matching the first spoken data with a pre-configured regular expression library using a regular matching module located in the game client corresponding to the target game, obtaining a first matching result, and transmitting the first matching result to the server. The information identification method according to feature 2.

4. The first model described above is obtained by training as follows: A first training set is obtained, the first training set includes a plurality of first training samples, the first training samples include first pre-configured account information and a first label corresponding to the first pre-configured account information, the first label is used to indicate that the first pre-configured account information is abnormal information, Based on the first training set described above, an initial model is trained, and an intermediate model is obtained. A second training set is obtained, the second training set includes multiple second training samples, and the second training samples include second pre-configured account information. The second training sample in the second training set is input to the intermediate model, the model output result is obtained, a second label is set for the second training sample based on the model output result, and the second label is used to indicate the probability that the second pre-set account information in the second training sample is abnormal information. The intermediate model is trained and the first model is obtained based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set. The information identification method according to feature 1.

5. Training the intermediate model and obtaining the first model based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set is: From the second training set, a second training sample is selected in which the second label is greater than a preset value, and the selected second training sample and the second label corresponding to the selected second training sample are set as the third training sample. This includes training the intermediate model based on a third training set and the first training set, and obtaining the first model. The information identification method according to feature 4.

6. The aforementioned initial model includes multiple sub-models, Training an initial model based on the first training set and obtaining an intermediate model includes training each of the submodels based on the first training set and obtaining a plurality of intermediate submodels. Inputting the second training sample from the second training set into the intermediate model, obtaining the model output result, and setting a second label for the second training sample based on the model output result includes inputting the second training sample from the second training set into each of the intermediate submodels, obtaining the output result for each of the intermediate submodels, and determining a second label for the second training sample based on the output result for each of the intermediate submodels. Training the intermediate model and obtaining the first model based on the first training set, the second training set, and a second label corresponding to the second training sample in the second training set includes training each of the intermediate submodels based on the first training set, the second training set, and a second label corresponding to the second training sample in the second training set, obtaining a plurality of trained intermediate submodels, and determining the plurality of trained intermediate submodels as the first model. The information identification method according to feature 5.

7. Determining the second label of the second training sample based on the output results of each intermediate submodel is: The average value of the output results of each intermediate submodel is determined as the second label of the second training sample. The process involves determining a second label for the second training sample based on the output results of each intermediate submodel and the output weights corresponding to each intermediate submodel, wherein the output weights corresponding to each intermediate submodel are determined based on second pre-configured account information in the second training sample. The information identification method according to feature 6.

8. The first model described above includes a number of trained intermediate submodels, Inputting the account information of the first game account in the aforementioned pre-configured database into a pre-trained first model and obtaining a first output result is: The account information of the first game account in the pre-configured database is input to each of the trained intermediate submodels, and the output result corresponding to each trained intermediate submodel is obtained. The first output is determined by determining the average value of the output results corresponding to each trained intermediate submodel as the first output, or by determining the first output based on the output results corresponding to each trained intermediate submodel and the output weights corresponding to each of the trained intermediate submodels, wherein the output weights corresponding to the trained intermediate submodels are determined based on information contained in the account information of the first game account. The information identification method according to feature 1.

9. The second model described above is obtained by training as follows: A fourth training set is obtained, the fourth training set includes a plurality of fourth training samples, the fourth training sample includes pre-configured speech data and a fourth label corresponding to the pre-configured speech data, the fourth label is used to indicate whether the pre-configured speech data is an abnormal speech, and if the pre-configured speech is an abnormal speech, it includes at least one of the following information: abnormal information, contact information, and user identity information, corresponding to the pre-configured speech data. Based on the fourth training set and a pre-configured dynamic learning rate scheduling policy, the second model is trained, and the trained second model is obtained. The information identification method according to feature 1.

10. The pre-configured speech data included in the fourth training sample is obtained by dividing the original speech data using a sliding window, and when dividing the original speech data, there is an overlapping portion between the preceding and succeeding speech data. The information identification method according to feature 9.

11. This involves performing fine-tuning on the trained second model and obtaining the fine-tuned second model. This further includes compressing and combining the finely tuned second model to obtain the final second model, and then using the final second model to predict speech data. The information identification method according to feature 9.

12. An information identification device comprising a data acquisition module, an information storage module, a first identification module, and a second identification module, The data acquisition module is configured to acquire first statement data corresponding to a first game account in the target game, perform abnormal content identification processing on the first statement data, and acquire the identification result. The information storage module is configured such that, if the identification result indicates that the first statement data contains abnormal content, it stores the first statement data in a pre-configured database, retrieves the account information of the first game account, and stores the account information of the first game account in association with the first statement data in the pre-configured database. The first identification module is configured to input the account information of the first game account in the pre-configured database into a pre-trained first model and obtain a first output result, the first output result being used to indicate the probability that the first game account is an abnormal account. The second identification module is configured to input the first utterance data from the pre-configured database into a pre-trained second model and obtain a second output result if the probability corresponding to the first output result is greater than a pre-configured probability threshold, the second output result includes at least one of the following: information indicating that the first utterance data is normal, abnormal information corresponding to the first utterance data, contact information, and user identity information. An information identification device characterized by the following features.

13. The system includes a processor and memory, the memory storing device-executable instructions that can be executed by the processor, and the processor executes the device-executable instructions to realize the information identification method described in any one of claims 1 to 11. An electronic device characterized by the following features.

14. A computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the information identification method described in any one of claims 1 to 11. A computer-readable storage medium characterized by the following features.

Citation Information

Patent Citations

  • Game data monitoring system and game data monitoring method

    CN113318454A

  • Method and device for identifying abnormal account in game, electronic equipment and storage medium

    CN113440856A

  • SYSTEM AND METHOD FOR VERIFYING GAME PLAY-RELATED ACTIVITY - Patent application

    JP2023545108A

  • Automatic classification and reporting of inappropriate language in online applications

    US20210370188A1

  • Offensive chat filtering using machine learning models

    US20220284884A1