Information identification method, device, and electronic device

The method and device enhance gaming safety by identifying abnormal accounts and extracting relevant information through a two-stage model, addressing illegal activities and privacy concerns in gaming platforms.

JP2025174795AActive Publication Date: 2025-11-28NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024133754
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-17
Filing Date
2024-08-09
Publication Date
2025-11-28
Estimated Expiration
2044-08-09

AI Technical Summary

Technical Problem

Illegal activities involving minors on third-party gaming platforms compromise game credibility and user privacy, making it difficult for game companies to gather evidence due to strict privacy regulations overseas.

Method used

An information identification method and device that performs risk assessment and anomaly detection using user behavioral data and comment content, involving a two-stage model process to identify abnormal accounts and extract relevant information.

Benefits of technology

Provides a safer gaming environment by efficiently identifying and mitigating illegal activities, enhancing user experience and compliance with privacy regulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025174795000001_ABST
    Figure 2025174795000001_ABST
Patent Text Reader

Abstract

To provide an information identification method, a device, and an electronic device for improving a user's game experience by providing the user with a safe game environment by determining an abnormal game account and remark in a game.SOLUTION: A method acquires an identification result by acquiring first remark data corresponding to a first game account, acquires account information of the first game account when it is shown that the remark data include abnormal contents, stores it in a preset database in association with the first remark data, inputs account information of the first game account to a first model, acquires a possibility of an abnormal account, inputs the first remark data in a preset database to a second model when the probability is larger than a preset probability threshold, and acquires at least one piece of information from among normality of a remark, abnormality information corresponding to the first remark data, and contact information and identity information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the field of data processing technology, and in particular to an information identification method, device and electronic equipment. [Background technology]

[0002] In some popular games, some illegal actors have been seen luring minors to third-party platforms to carry out illegal activities, which has had a negative impact on the games, seriously damaging the games' credibility and causing users to lose interest. However, overseas markets have strict requirements for privacy and personal information protection, and overseas game companies cannot obtain users' personal information or freely examine users' speech records. Secondly, illegal actors often luring minors to third-party platforms, which makes it difficult to leave behind illegal evidence, and game companies have no sufficient clues to investigate. Summary of the Invention

[0003] The object of the present invention is to provide an information identification method, device, and electronic device that can provide a safe gaming environment for users and improve their gaming experience by performing risk assessment and anomaly identification using user behavioral data and comment content.

[0004] According to a first aspect, the present invention provides an information identification method, the method including: obtaining first utterance data corresponding to a first game account in a target game; performing an anomalous content identification process on the first utterance data to obtain an identification result; if the identification result indicates that the first utterance data contains anomalous content, storing the first utterance data in a predetermined database, obtaining account information of the first game account, and associating the account information of the first game account with the first utterance data and storing them in the predetermined database; inputting the account information of the first game account in the predetermined database into a pre-trained first model to obtain a first output result, wherein the first output result is used to indicate a probability that the first game account is an anomalous account; if the probability corresponding to the first output result is greater than a predetermined probability threshold, inputting the first utterance data in the predetermined database into a pre-trained second model to obtain a second output result, wherein the second output result includes at least one of information that the utterance is normal, anomalous information corresponding to the first utterance data, contact information, and user identity information.

[0005] According to a second aspect, the present invention provides an information identification device, the device comprising: a data acquisition module, an information storage module, a first identification module, and a second identification module, wherein the data acquisition module is configured to acquire first utterance data corresponding to a first game account in a target game, perform an abnormal content identification process on the first utterance data, and acquire an identification result, and the information storage module is configured to, when the identification result indicates that the first utterance data contains abnormal content, store the first utterance data in a preset database, acquire account information of the first game account, and associate the account information of the first game account with the first utterance data and store them in the preset database. The first identification module is configured to input account information of the first game account in the predetermined database into a pre-trained first model to obtain a first output result, the first output result being used to indicate the probability that the first game account is an abnormal account; and the second identification module is configured to input the first utterance data in the predetermined database into a pre-trained second model to obtain a second output result if the probability corresponding to the first output result is greater than a predetermined probability threshold, the second output result including at least one of information that the utterance is normal, abnormal information corresponding to the first utterance data, contact information, and user identity information.

[0006] According to a third aspect, the present invention provides an electronic device including a processor and a memory, wherein the memory stores machine-executable instructions executable by the processor, and the processor executes the machine-executable instructions to realize the above-mentioned information identification method.

[0007] According to a fourth aspect, the present invention provides a computer-readable storage medium having stored thereon computer-executable instructions which, when called upon and executed by a processor, cause the processor to implement the information identification method set forth above.

[0008] Embodiments of the present invention provide the following beneficial effects.

[0009] The information identification method, device, and electronic device of the present invention first obtain first utterance data corresponding to a first game account in a target game, perform an anomalous content identification process on the first utterance data, and obtain an identification result. If the identification result indicates that the first utterance data contains anomalous content, store the first utterance data in a predetermined database, obtain account information for the first game account, and associate the account information of the first game account with the first utterance data and store them in the predetermined database. Next, input the account information of the first game account in the predetermined database into a pre-trained first model and obtain a first output result. The first output result is used to indicate the probability that the first game account is an anomalous account. If the probability corresponding to the first output result is greater than a predetermined probability threshold, input the first utterance data in the predetermined database into a pre-trained second model and obtain a second output result. The second output result includes at least one of information indicating that the utterance is normal, anomalous information corresponding to the first utterance data, contact information, and user identity information. This method utilizes the account information and speech data of the game account to perform risk assessment and anomaly identification for the game account, providing a safe gaming environment for users and helping to improve the user's gaming experience.

[0010] Additional features and advantages of the invention will be set forth in the following specification, or some features and advantages may be inferred or determined inherently from the specification, or may be learned by practicing the above-described techniques of the invention.

[0011] In order to make the above objects, features and advantages of the present invention more apparent, preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. [Brief explanation of the drawings]

[0012] In order to more clearly describe the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings that need to be used in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention, and those skilled in the art can also obtain other drawings based on these drawings without any creative efforts. [Figure 1] 2 is a flowchart of an information identification method according to an embodiment of the present invention. [Figure 2] 4 is a flowchart of information identification according to an embodiment of the present invention. [Figure 3] 1 is a structural schematic diagram of an information identification device according to an embodiment of the present invention; [Figure 4] 1 is a structural schematic diagram of an electronic device according to an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0013] In order to make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be described clearly and completely below with reference to the drawings in the embodiments of the present invention, and it is obvious that the described embodiments are only some embodiments of the present invention, and not all embodiments. Generally, the components in the embodiments of the present invention described and illustrated in the drawings herein can be arranged and designed in a variety of different configurations.

[0014] Therefore, the detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claims, but merely illustrates specific embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative labor also fall within the protection scope of the present invention.

[0015] In some popular games, some illegal actors have been seen luring minors to third-party platforms to engage in illegal activities, which has had a negative impact on the games, seriously damaging the games' credibility and causing users to lose interest. However, overseas markets have strict requirements for privacy and personal information protection, and overseas game companies cannot obtain users' personal information or freely examine users' speech records. Therefore, illegal actors often luring minors to third-party platforms, which makes it difficult to leave behind illegal evidence, and game companies have no sufficient clues to investigate.

[0016] Based on the above problem, the embodiments of the present invention provide an information identification method, device and electronic device, which can be applied to any game information identification scenario, especially to the information identification scenario for fraudulent game accounts.

[0017] To facilitate understanding of the embodiments of the present invention, the information identification method disclosed in the embodiments of the present invention will be described in detail first. As shown in FIG. 1, the method includes the following specific steps:

[0018] In step S102, first utterance data corresponding to a first game account in the target game is obtained, and abnormal content identification processing is performed on the first utterance data to obtain an identification result.

[0019] In a specific implementation, the target game may be any game, the first game account may be any game account registered in the target game, and the first utterance data corresponding to the first game account includes all utterance data sent by the user in the target game through the first game account, including game comments, conversations within the game, or game messages, etc.

[0020] In practical application, when a user sends utterance data in a target game through a first game account, the utterance data is immediately acquired, an abnormal content identification process is performed on the utterance data, and an identification result corresponding to the utterance data is obtained, and the identification result is used to indicate whether the utterance data contains abnormal content or not. Specifically, when identifying the abnormal content of the utterance data, the utterance data is matched with a keyword in a preset keyword library, and if the utterance data matches a keyword in the keyword library, the identification result may indicate that the utterance data contains abnormal content, or the utterance data may be matched with a preset regular expression, and if both match, the identification result indicates that the utterance data contains abnormal content.

[0021] In step S104, the identification result indicates that the first utterance data contains abnormal content, the first utterance data is stored in a preset database, account information of the first game account is obtained, and the account information of the first game account is associated with the first utterance data and stored in the preset database.

[0022] In a specific implementation, if the identification result indicates that the first utterance data does not contain anomalous content, no processing is performed on the first utterance data. If the identification result indicates that the first utterance data does contain anomalous content, the first utterance data is not stored in the preset database, and preparation is made for further artificial screening and model analysis processing. In addition to storing the first utterance data in the preset database, it is necessary to determine whether the preset database contains account information for a first game account corresponding to the first utterance data. If the account information for the first game account is contained, it is explained that the preset database contains utterance data corresponding to the first game account, and the first utterance data and the account information for the first game account are directly associated and stored in the preset database. If the account information for the first game account is not contained in the preset database, it is necessary to obtain account information for the first game account from the target game, and associate the account information for the first game account with the first utterance data and store it in the preset database.

[0023] Specifically, the pre-set database may be a MySQL database, a relational database, a DB database, etc., and may be specifically determined according to development needs. The account information of the first game account includes at least one of information such as the number of messages in the first game account's friend channel, the number of messages in stranger's private message channel, the number of messages in the team channel, the number of messages in the room channel, the number of messages in other channels, the number of matches, the number of masked word hits, the number of people added to the blacklist, the number of people added to the blacklist, the number of friend requests, the number of friend deletions, the number of deleted friends, the number of people moved from the blacklist, the number of friend requests accepted, the number of times the friend requests were terminated, the time the account was created, and the number of IP addresses used by the account.

[0024] In step S106, the account information of the first game account in the preset database is input into a pre-trained first model to obtain a first output result, where the first output result is used to indicate the probability that the first game account is an abnormal account.

[0025] If the account information of the first game account is stored in a pre-configured database, the account information of the first game account needs to be input into a pre-trained first model, which can capture the behavioral characteristics of the user who uses the first game account based on the account information of the first game account, and thereby predict the probability that the first game account is an abnormal account. The training process of the first model will be described in the following examples.

[0026] In step S108, if the probability corresponding to the first output result is greater than a preset probability threshold, the first utterance data in the preset database is input into a pre-trained second model to obtain a second output result, where the second output result includes at least one of information that the utterance is normal, abnormal information corresponding to the first utterance data, contact information, and user identity information.

[0027] In a specific implementation, the preset probability threshold may be determined according to development needs, for example, the preset probability threshold may be set to 80% or 90%, etc. If the probability corresponding to the first output result of the first model is equal to or less than the preset probability threshold, the first game account is declared to be a normal game account, and no processing is performed on the first game account; if the probability corresponding to the first output result is greater than the preset probability threshold, the first game account is declared to be an abnormal account, so that all utterance data corresponding to the first game account in the preset database needs to be input into a second model, and the second model performs natural language identification on the input utterance data to extract text features from the input utterance data and determine whether the input utterance data is an abnormal utterance according to the text features; if the input utterance data is an abnormal utterance, obtain one or more of anomaly information, contact information, and user identity information from the input utterance data.

[0028] The user's identity information can be used to determine whether the user using the first game account is a minor or whether the speech data contains information related to a minor. In practical applications, abnormal content in the speech data is determined based on the second output result, and the abnormal content is masked or otherwise processed, thereby providing a safer language environment and gaming environment for minors and ensuring their safety in games.

[0029] An embodiment of the present invention provides an information identification method that uses account information and comment data of a game account to perform risk assessment and anomaly identification for the game account, thereby providing a safe gaming environment for users and helping to improve users' gaming experience.

[0030] The following example is intended to explain a method for identifying anomalies in utterance data.

[0031] Specifically, the specific process of performing the abnormal content identification process on the first utterance data and acquiring the identification result includes at least one of the following.

[0032] First, the first utterance data is matched with a predetermined regular expression library to obtain a first matching result, and the first matching result is determined as an identification result corresponding to the first utterance data, where if the first utterance data is matched with a regular expression in the regular expression library, the first matching result indicates that the first utterance data contains abnormal content.

[0033] In a specific implementation, the regular expressions in the regular expression library are pre-defined according to development needs, and the keywords that each regular expression can match can be varied in many ways. For example, the regular expression " / live.*? / " can match keyword variations such as "where do you live" and "where do you live", thereby broadening the scope of keyword matching using regular expressions.

[0034] Second, the first utterance data is matched with a preset keyword library to obtain a second matching result, and the second matching result is determined as an identification result corresponding to the first utterance data, where if the first utterance data hits a keyword in the keyword library, the second matching result indicates that the first utterance data contains abnormal content.

[0035] The keyword library includes many keywords, which may be determined according to development needs, for example, the keywords may be "sex," "where to go," etc. In a specific implementation, if the first utterance data includes a keyword in the keyword library, or if the first utterance data includes a predetermined number of keywords, it is determined that the first utterance data matches a keyword in the keyword library.

[0036] In one specific embodiment, in order to cover more scenes and speed up matching without affecting the user experience, the present invention can simultaneously perform anomaly matching using a regular expression library and a keyword library. Figure 2 is a flowchart of information identification according to an embodiment of the present invention. The regular keyword library matching process in Figure 2 includes: when a user submits utterance data in a target game, the utterance data simultaneously passes through a regular matching module and a keyword search module in an SDK (Software Development Kit), where the regular matching module matches the utterance data using a regular expression library, and the keyword search module matches the utterance data using a keyword library, where the regular matching module can increase keyword coverage and the keyword search module can be responsible for quickly searching for anomalous content, so that the method can quickly and efficiently block known risk signals (the risk signals correspond to utterance data containing anomalous content). Finally, the suspected anomalous utterance data is stored in a pre-defined database for further manual screening and model analysis, where the suspected anomalous utterance data is utterance data matched with a regular expression in the regular expression library or utterance data that hits a keyword in the keyword library.

[0037] Specifically, the keyword search algorithm has low complexity (O(1)), can quickly match known abnormal content, and directly searches tables without requiring complex pattern matching, resulting in extremely fast search speeds and effective coverage of abnormal scenes with clear intent. For example, the keyword "sex" can match specific abnormal scenes. Relying solely on a keyword library is insufficient because it is difficult to cover all variants and implicit abnormal terms. Therefore, the present invention introduces a regular expression library to compensate for the deficiencies of keyword libraries.

[0038] In an alternative embodiment, the method provided by the present invention is used in a server, but the regular matching module of the SDK can be located in the game client and the keyword module can be located in the server to perform the search so that the user cannot detect the search matching. Therefore, the regular matching module located in the game client corresponding to the target game matches the first utterance data with a pre-defined regular expression library to obtain a first matching result, and then transmits the first matching result to the server. The regular matching module can also intercept messages that the user clicked to send in the target game but were not sent, thereby obtaining more utterance data.

[0039] The above method can quickly and accurately identify whether the comment data contains abnormal content, and then further determine the game account corresponding to the comment data containing abnormal content to determine whether the game account is an abnormal account.

[0040] The following example is used to explain the training scheme of the first model.

[0041] Specifically, the first model is obtained by training in steps 10 to 14 below.

[0042] In step 10, a first training set is obtained, the first training set includes a plurality of first training samples, the first training samples include first preset account information and a first label corresponding to the first preset account information, and the first label is used to indicate that the first preset account information is abnormal information.

[0043] In a specific implementation, the first training sample in the first training set is from user data that has been accused and verified, that is, the user data is account information corresponding to a game account that has been accused and verified of abnormal behavior, and the user data corresponds to the first preset account information. Since the user data is the account information of a game account in which abnormal behavior exists, a first label corresponding to the first preset account information is used to indicate that the first preset account information is abnormal information.

[0044] In step 11, an initial model is trained based on the first training set to obtain an intermediate model.

[0045] In a specific implementation, a target first training sample can be first selected from a first training set, the first preset account information of the target first training sample is input into an initial model to obtain a model output result, and then the model parameters of the initial model are adjusted according to the model output result and the first label of the target first training sample to obtain an initial model after parameter adjustment, and the step of selecting a target first training sample from the first training set is continued until the initial model after parameter adjustment converges, and the converged initial model is determined as an intermediate model, where each first training sample in the first training set can participate in model training as a single target first training sample.

[0046] In step 12, a second training set is obtained, where the second training set includes a plurality of second training samples, and the second training samples include second preset account information.

[0047] The second training sample included in the second training set is second preset account information to which no label is set, and the second preset account information may be account information of any game account acquired from the target game. In other words, the second preset account information is account information that has not been verified as being an abnormal account, and therefore a label cannot be automatically set for the second preset account information.

[0048] In step 13, a second training sample in the second training set is input into the intermediate model to obtain a model output result, and a second label is set for the second training sample based on the model output result, and the second label is used to indicate the probability that the second preset account information in the second training sample is anomalous information.

[0049] In a specific implementation, since the amount of user data that has been accused and verified is small, a high-quality model can be constructed that can well capture the essential principles of user behavioral features and has good generalization ability to unseen data, so that the acquired behavioral feature data can be fully utilized to set labels in the second training set using pseudo-labeling technology.

[0050] In practical application, the second training sample in the second training set is input to the intermediate model to obtain the model output result, and then the model output result can be used as the second label of the second training sample, or the product of the model output result and the preset weight can be used as the second label of the second training sample.

[0051] In step 14, an intermediate model is trained based on the first training set, the second training set and the second label corresponding to the second training sample in the second training set to obtain a first model.

[0052] In a specific implementation, an intermediate model can be trained based on all first training samples in the first training set, some or all second training samples in the second training set, and second labels corresponding to the second training samples to obtain the first model.

[0053] In an alternative embodiment, a second training sample whose second label is greater than a preset value is selected from the second training set, and the selected second training sample and the second label corresponding to the selected second training sample are used as a third training sample. An intermediate model is trained based on the third training set and the first training set to obtain the first model. The preset value may be determined according to development needs, for example, the preset value may be 80% or 90%.

[0054] In another alternative embodiment, the second training samples may be sorted in descending order of probability of corresponding to the second label, and a specified number of the top second training samples and the second labels corresponding to the second training samples may be selected as the third training samples.

[0055] In one specific embodiment, the initial model may be a machine learning model or a neural network model, etc. For example, the initial model may be an XGBoost model, where the XGBoost model has parallelization characteristics and can generate an excellent model quickly and efficiently.

[0056] In another alternative embodiment, a single initial model may overfit the training data, thereby reducing its adaptability to new, unseen data. Therefore, the initial model in the present invention may include multiple submodels, whereby each submodel is trained based on a first training set to obtain multiple intermediate submodels, and second training samples in a second training set are input to each intermediate submodel to obtain output results of each intermediate submodel, and second labels of the second training samples are determined based on the output results of each intermediate submodel. Then, each intermediate submodel is trained based on the first training set, the second training set, and the second labels corresponding to the second training samples in the second training set to obtain multiple trained intermediate submodels, and the trained multiple intermediate submodels are determined as the first model. Here, the training methods of each submodel in the initial model are the same, and different submodels are trained separately.

[0057] In a specific implementation, the specific process of determining the second label of the second training sample according to the output results of each intermediate sub-model may include any of the following:

[0058] First, the average value of the output results of each intermediate sub-model is determined as the second label of the second training sample. For example, if the number of intermediate sub-models is three, the second label of the second training sample is the sum of the output results of the three intermediate sub-models divided by three.

[0059] Second, determine a second label of a second training sample based on the output result of each intermediate sub-model and the output weight corresponding to each intermediate sub-model, where the output weight corresponding to the intermediate sub-model is determined based on second preset account information in the second training sample.

[0060] In a specific implementation, the information content corresponding to the second preset account information is different, and the output weight corresponding to each intermediate sub-model is also different, and which information content in the second preset account information affects the output weight corresponding to the intermediate sub-model is preset. For example, if the number of intermediate sub-models is three and the number of comments on the friend channel in the second preset account information is large, the output weight corresponding to the first intermediate sub-model is high, and the output weights corresponding to the other two intermediate sub-models are low; if the number of deleted friends is large, the output weight corresponding to the second intermediate sub-model is high, etc.

[0061] In one specific embodiment, the initial model shown in FIG. 2 includes three boosting models, XGBoost, CatBoost, and LightGBM. These three sub-models are trained using a first training set to obtain intermediate sub-models corresponding to the three sub-models. These three intermediate sub-models are then first applied to a number of unlabeled data samples (corresponding to the second training set) to perform prediction and pseudo-labeling. Next, from these pseudo-labeled data (corresponding to the second training samples and the second labels corresponding to the second training samples), the parts with the highest model prediction probability values ​​are selected and considered as additional labeled samples (corresponding to the third training set) and added to the first training set. These three intermediate sub-models are then trained based on the third training set and the first training set to obtain three trained intermediate sub-models.

[0062] To further improve the robustness of the model, the large amount of sample data augmented by pseudo-labeling techniques can be trained in parallel with three different boosting tree models, XGBoost, CatBoost, and LightGBM, using a 5-fold cross-validation (5-fold CV) method. Finally, the output probabilities of the three models in the 5-fold validation are stacked, i.e., the output probabilities of the three models are used as new features, and these probability features are integrated to obtain the final predicted probability (corresponding to the second label above).

[0063] Based on the above description, the pre-trained first model includes a plurality of trained intermediate sub-models, and a specific process of inputting the account information of the first game account in the above-mentioned preset database into the pre-trained first model and obtaining a first output result may include: inputting the account information of the first game account in the preset database into a plurality of trained intermediate sub-models respectively, and obtaining an output result corresponding to each trained intermediate sub-model; and determining the average value of the output results corresponding to each trained intermediate sub-model as the first output result, or determining the first output result based on the output result corresponding to each trained intermediate sub-model and the output weight corresponding to each trained intermediate sub-model, wherein the output weight corresponding to the trained intermediate sub-model is determined based on information included in the account information of the first game account.

[0064] The above self-training method successfully expands the size and diversity of the training set, making the data distribution seen by the model during training more similar to the real situation, thereby effectively improving the model's generalization and adaptability to unseen data.

[0065] The following example is used to illustrate the training method of the second model.

[0066] Specifically, the above-mentioned second model is obtained by training as follows: a fourth training set is obtained, the fourth training set includes a plurality of fourth training samples, the fourth training samples include predetermined utterance data and a fourth label corresponding to the predetermined utterance data, the fourth label is used to indicate whether the predetermined utterance data is an abnormal utterance, and if the predetermined utterance is an abnormal utterance, the fourth label includes at least one of abnormal information, contact information, and user identity information corresponding to the predetermined utterance data; and the second model is trained based on the fourth training set and the predetermined dynamic learning rate scheduling policy, and a trained second model is obtained.

[0067] In a specific implementation, the preset utterance data in the fourth training samples in the fourth training set is utterance data in the game corresponding to the user data accused of detecting an abnormality, and the number of training samples included in the fourth training set may be determined according to development needs, for example, the number of training samples may be 1000 or 800. After obtaining the preset utterance data in the fourth training samples, it is necessary to manually label the preset utterance data to obtain a fourth label corresponding to the preset utterance data, where the fourth label is used to indicate whether the preset utterance data is an abnormal utterance, and if the preset utterance is an abnormal utterance, it includes at least one of the anomaly information, contact information, and user identity information corresponding to the preset utterance data.

[0068] In one specific example, manual labeling can be performed on the pre-set utterance data in four ways: by specifying the user's identity, by specifying abnormal scenes, and by including contact information and normal content (corresponding to the utterance being normal). Because context often needs to be combined to accurately understand the meaning of certain words or phrases in the pre-set utterance data, context information is also tightly integrated into the labeling process. For example, for the Japanese phrase "Yaru?" (do you play?) the fourth label assigned to the labeling result will differ depending on the context.

[0069] In a specific implementation, the second model may be optimized and adjusted based on LORA to improve the convergence speed and fitting ability during the model training process. First, instead of using a uniform global learning rate, the learning rate parameters of each layer of the network in the second model can be individually set based on LORA technology. This allows different layers to adopt different learning rates according to their own characteristics, thereby making network parameter updates more precise and purposeful. Second, the present invention introduces a dynamic learning rate scheduling policy to monitor changes in the loss function during the training process in real time and automatically adjust the learning rate of each layer. When the loss function tends to slow down, the learning rate is appropriately increased to speed up convergence, and when the loss function fluctuates significantly, the learning rate is decreased to avoid oscillations and divergences. This dynamic adjustment mechanism contributes to capturing the optimal convergence point and significantly improves the convergence efficiency of the second model within a limited number of training steps.

[0070] In specific implementation, the second model may be a large-scale language model (also called an LLM model) or other neural network model, and may be determined according to specific development needs. Here, the present invention may adopt the Llama2 large-scale model in the LLM model, which is a general-purpose large-scale model that can distinguish 150 languages ​​such as Japanese, Chinese, and English. During the training process, three pieces of data need to be input: Prompt is a prompt word, and its content is "You are a valid language assistant and can identify abnormal utterances. Please judge based on the following data: label is normal, user identity information, contact information, abnormal information, user identity information + abnormal information, user identity information + contact information, user identity information + abnormal information + contact information, abnormal information + contact information", and text is "Hello, I'm xxx, would you like to be friends with me? I'm a 16-year-old student, where are you from? Hey, let's play together?" and output is "user identity information + contact information", where the user identity information here may be the identity of a user whose age is below a preset age threshold, and the preset age threshold may be determined according to development needs.

[0071] In an optional embodiment, the preset utterance data included in the fourth training sample is obtained by dividing the original utterance data using a sliding window, and when dividing the original utterance data, there is an overlapping portion between the divided utterance data in the front and rear stages.

[0072] In a specific implementation, the present invention performs a rational segmentation process on labeled training samples. Experiments have shown that if a text segment (corresponding to the preset utterance data) is too long or too short, it will affect the accurate identification and understanding of the text content by the large-scale language model (corresponding to the second model). Therefore, the present invention uses a sliding window method to segment the preset utterance data, and appropriately adjusts the length of the sliding window to maintain sufficient text length while expanding the context view as much as possible, allowing the second model to better capture contextual information and improve its learning ability for the text content. Next, to avoid misunderstandings caused by context loss, an overlap mechanism is introduced, i.e., by maintaining a certain overlap area between two adjacent sliding windows, it is possible to ensure that the divided utterance data in the front and back sections have overlapping portions, and important contextual information is simultaneously present in the adjacent windows.

[0073] In an optional embodiment, a compression and combination process can be performed on the second model so that the second model can be smoothly deployed on the server. Specifically, a fine-tuning process is performed on the trained second model to obtain the fine-tuned second model, a compression and combination process is performed on the fine-tuned second model to obtain a final second model, and the final second model can be used to predict utterance data.

[0074] In a specific implementation, the fine-tuning process of conventional LORA (Low-Rank Adaptation) only updates a small amount of offset and attention parameters in the second model, leaving most of the parameters in the pre-trained second model unchanged. While this method can accelerate training and save memory, it requires loading the full pre-trained model parameters in the inference stage, resulting in significant storage and computation overhead. Therefore, QLORA (Quantized Low-Rank Adaptation) technology can be used to compress and combine LORA fine-tuned parameters to generate dedicated "inference weights." The specific process is as follows:

[0075] 1. Derive and calculate the equivalent full parameter matrix based on the updated small offset and attention parameters of LORA.

[0076] 2. Compress these full parameter matrices, e.g., by techniques such as matrix decomposition and quantization to reduce the parameter dimension and memory occupation.

[0077] 3. Combine the compressed parameters with the raw parameters of a second pre-trained model to generate a new inference weight tensor.

[0078] The above method comprehensively utilizes user account information and speech data to perform risk assessment and anomaly identification for users. Through intelligent modeling and big data analysis, suspected anomalies can be efficiently detected. To meet the specific needs of minor protection, the method combines a small labeling corpus and large-scale language modeling technology to successfully train a highly accurate multi-label language classification model, significantly reducing human resources. This model can accurately identify inappropriate content in various languages, providing a safer language environment for minors.

[0079] Corresponding to the above method embodiment, the embodiment of the present invention further provides an information identification device, and as shown in FIG. 3, the device includes: a data acquisition module 30, an information storage module 31, a first identification module 32, and a second identification module 33.

[0080] the data acquisition module 30 is configured to acquire first utterance data corresponding to a first game account in the target game, perform an abnormal content identification process on the first utterance data, and acquire an identification result; the information storage module 31 is configured, when the identification result indicates that the first utterance data contains abnormal content, to store the first utterance data in a preset database, obtain account information of the first game account, and associate the account information of the first game account with the first utterance data and store it in the preset database; The first identification module 32 is configured to input account information of the first game account in the preset database into a pre-trained first model to obtain a first output result, which is used to indicate a probability that the first game account is an abnormal account; The second identification module 33 is configured to input the first utterance data in the preset database into a pre-trained second model to obtain a second output result when the probability corresponding to the first output result is greater than a preset probability threshold, and the second output result includes at least one information of the normality of the utterance, abnormality information corresponding to the first utterance data, contact information, and user identity information.

[0081] According to the above information identification device, the method uses the account information and comment data of the game account to perform risk assessment and anomaly identification for the game account, thereby providing a safe gaming environment for users and helping to improve the user's gaming experience.

[0082] Specifically, the data acquisition module 30 is for matching first utterance data with a predetermined regular expression library, obtaining a first matching result, and determining the first matching result as an identification result corresponding to the first utterance data; if the first utterance data is matched with a regular expression in the regular expression library, the first matching result indicates that the first utterance data contains abnormal content; and / or, for matching the first utterance data with a predetermined keyword library, obtaining a second matching result, and determining the second matching result as a recognition result corresponding to the first utterance data; if the first utterance data hits a keyword in the keyword library, the second matching result indicates that the first utterance data contains abnormal content.

[0083] In a specific implementation, the above method is applied to a server, and the data acquisition module 30 of the above device is further configured to match the first utterance data with a pre-set regular expression library through a regular matching module located in a game client corresponding to the target game, obtain a first matching result, and send the first matching result to the server.

[0084] Further, the above-mentioned apparatus further includes a first model training module for obtaining a first training set, wherein the first training set includes a plurality of first training samples, the first training samples including first preset account information and a first label corresponding to the first preset account information, and the first label is used to indicate that the first preset account information is abnormal information; training an initial model based on the first training set to obtain an intermediate model; and obtaining a second training set, wherein the second training set includes a plurality of second training samples. the second training samples include second preset account information; the second training samples in the second training set are input into an intermediate model to obtain a model output result; a second label is set for the second training sample based on the model output result, the second label is used to indicate a probability that the second preset account information in the second training sample is anomalous information; and the intermediate model is trained based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set to obtain the first model.

[0085] In a specific implementation, the first model training module is further configured to: select second training samples from the second training set, whose second labels are greater than a predetermined value; set the selected second training samples and the second labels corresponding to the selected second training samples as third training samples; and train an intermediate model based on the third training set and the first training set to obtain the first model.

[0086] In an alternative embodiment, the initial model includes a plurality of sub-models, and the first model training module is further configured to: train each sub-model based on a first training set to obtain a plurality of intermediate sub-models; input second training samples in a second training set into each intermediate sub-model; obtain an output result of each intermediate sub-model; determine second labels of the second training samples based on the output result of each intermediate sub-model; train each intermediate sub-model based on the first training set, the second training set, and the second labels corresponding to the second training samples in the second training set; obtain the trained plurality of intermediate sub-models; and determine the trained plurality of intermediate sub-models as the first model.

[0087] Furthermore, determining a second label of the second training sample based on the output result of each intermediate sub-model includes either determining an average value of the output result of each intermediate sub-model as the second label of the second training sample, or determining a second label of the second training sample based on the output result of each intermediate sub-model and output weights corresponding to each intermediate sub-model, wherein the output weights corresponding to the intermediate sub-models are determined based on second predetermined account information in the second training sample.

[0088] In a specific implementation, the first model includes a plurality of trained intermediate sub-models, and the first identification module 32 is configured to input the account information of the first game account in the preset database into the plurality of trained intermediate sub-models respectively, obtain output results corresponding to each trained intermediate sub-model, and determine the average value of the output results corresponding to each trained intermediate sub-model as the first output result, or determine the first output result based on the output results corresponding to each trained intermediate sub-model and the output weights corresponding to each trained intermediate sub-model, where the output weights corresponding to the trained intermediate sub-models are determined based on the information included in the account information of the first game account.

[0089] Further, the above-mentioned device further includes a second model training module for obtaining a fourth training set, wherein the fourth training set includes a plurality of fourth training samples, wherein the fourth training samples include predetermined utterance data and a fourth label corresponding to the predetermined utterance data, wherein the fourth label is used to indicate whether the predetermined utterance data is an abnormal utterance, and if the predetermined utterance is an abnormal utterance, includes at least one of anomaly information, contact information, and user identity information corresponding to the predetermined utterance data; and trains a second model based on the fourth training set and the predetermined dynamic learning rate scheduling policy, and obtains a trained second model.

[0090] In an optional embodiment, the preset utterance data included in the fourth training sample is obtained by dividing the original utterance data using a sliding window, and when dividing the original utterance data, there is an overlapping portion between the divided utterance data in the front and rear stages.

[0091] Furthermore, the above-mentioned device further includes a model compression module, configured to perform a fine-tuning process on the trained second model to obtain a fine-tuned second model, perform a compression and combination process on the fine-tuned second model to obtain a final second model, and use the final second model to predict utterance data.

[0092] The information identification device according to the embodiments of the present disclosure has the same implementation principle and generated technical effects as the above-mentioned method embodiments, and will be described briefly. Where no reference is made to the device embodiments, reference may be made to the corresponding content in the above-mentioned method embodiments.

[0093] An embodiment of the present invention further provides an electronic device, and as shown in FIG. 4, the electronic device includes a processor and a memory, the memory stores machine-executable instructions executable by the processor, and the processor executes the machine-executable instructions to realize the above-mentioned information identification method.

[0094] Specifically, the information identification method includes: acquiring first utterance data corresponding to a first game account in a target game; performing an abnormal content identification process on the first utterance data; and obtaining an identification result. If the identification result indicates that the first utterance data contains abnormal content, storing the first utterance data in a predetermined database, acquiring account information of the first game account, and storing the account information of the first game account in association with the first utterance data in the predetermined database. Inputting the account information of the first game account in the predetermined database into a pre-trained first model and obtaining a first output result, wherein the first output result is used to indicate the probability that the first game account is an abnormal account. If the probability corresponding to the first output result is greater than a predetermined probability threshold, inputting the first utterance data in the predetermined database into a pre-trained second model and obtaining a second output result, wherein the second output result includes at least one of information indicating that the utterance is normal, abnormal information corresponding to the first utterance data, contact information, and user identity information.

[0095] The above information identification method uses the account information and comment data of a game account to perform risk assessment and abnormality identification for the game account, thereby providing a safe gaming environment for users and helping to improve the user's gaming experience.

[0096] In an optional embodiment, performing an abnormal content identification process on the first utterance data and obtaining an identification result includes at least one of: matching the first utterance data with a predetermined regular expression library, obtaining a first matching result, and determining the first matching result as an identification result corresponding to the first utterance data, wherein if the first utterance data is matched with a regular expression in the regular expression library, the first matching result indicates that the first utterance data contains abnormal content; and matching the first utterance data with a predetermined keyword library, obtaining a second matching result, and determining the second matching result as a recognition result corresponding to the first utterance data, wherein if the first utterance data hits a keyword in the keyword library, the second matching result indicates that the first utterance data contains abnormal content.

[0097] In an alternative embodiment, the method is applied to a server, and the method further includes matching the first utterance data with a predetermined regular expression library by a regular matching module located in a game client corresponding to the target game, obtaining a first matching result, and sending the first matching result to the server.

[0098] In an alternative embodiment, the first model is obtained by training as follows: obtaining a first training set, the first training set including a plurality of first training samples, the first training samples including first preset account information and a first label corresponding to the first preset account information, the first label being used to indicate that the first preset account information is abnormal information; training an initial model based on the first training set to obtain an intermediate model; obtaining a second training set, the second training set including a plurality of second training samples; The second training sample includes a sample, and the second training sample includes second preset account information. The second training sample in the second training set is input into an intermediate model to obtain a model output result. A second label is set for the second training sample based on the model output result, and the second label is used to indicate a probability that the second preset account information in the second training sample is anomalous information. The intermediate model is trained based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set to obtain a first model.

[0099] In an optional embodiment, training an intermediate model based on the first training set, the second training set, and a second label corresponding to the second training sample in the second training set to obtain the first model includes: selecting a second training sample from the second training set whose second label is greater than a predetermined value, and setting the selected second training sample and the second label corresponding to the selected second training sample as a third training sample; and training an intermediate model based on the third training set and the first training set to obtain the first model.

[0100] In an alternative embodiment, the initial model includes a plurality of sub-models, and training the initial model based on the first training set to obtain intermediate models includes training each sub-model based on the first training set to obtain a plurality of intermediate sub-models, inputting second training samples in a second training set into the intermediate models and obtaining model output results, and setting second labels for the second training samples based on the model output results includes inputting the second training samples in the second training set into each intermediate sub-model, obtaining output results for each intermediate sub-model, and setting second labels for each intermediate sub-model. determining a second label of the second training sample based on an output result of the neural network model; training an intermediate model based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set to obtain the first model; training each intermediate sub-model based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set to obtain a plurality of trained intermediate sub-models; and determining the trained plurality of intermediate sub-models as the first model.

[0101] In an optional embodiment, determining a second label of the second training sample based on the output results of each intermediate sub-model includes either determining an average value of the output results of each intermediate sub-model as the second label of the second training sample, or determining a second label of the second training sample based on the output results of each intermediate sub-model and output weights corresponding to each intermediate sub-model, wherein the output weights corresponding to the intermediate sub-models are determined based on second predetermined account information in the second training sample.

[0102] In an alternative embodiment, the first model includes a plurality of trained intermediate sub-models, and inputting the account information of the first game account in the predetermined database into the pre-trained first model and obtaining a first output result includes inputting the account information of the first game account in the predetermined database into the plurality of trained intermediate sub-models, respectively, and obtaining an output result corresponding to each trained intermediate sub-model; and determining an average value of the output results corresponding to each trained intermediate sub-model as the first output result, or determining the first output result based on the output results corresponding to each trained intermediate sub-model and output weights corresponding to each trained intermediate sub-model, wherein the output weights corresponding to the trained intermediate sub-models are determined based on information included in the account information of the first game account.

[0103] In an optional embodiment, the second model is obtained by training as follows: a fourth training set is obtained, the fourth training set includes a plurality of fourth training samples, the fourth training samples include predetermined utterance data and a fourth label corresponding to the predetermined utterance data, the fourth label is used to indicate whether the predetermined utterance data is an abnormal utterance, and if the predetermined utterance is an abnormal utterance, the fourth label includes at least one of anomaly information, contact information, and user identity information corresponding to the predetermined utterance data; and the second model is trained based on the fourth training set and a predetermined dynamic learning rate scheduling policy, and a trained second model is obtained.

[0104] In an optional embodiment, the preset utterance data included in the fourth training sample is obtained by dividing the original utterance data using a sliding window, and when dividing the original utterance data, there is an overlapping portion between the divided utterance data in the front and rear stages.

[0105] In an alternative embodiment, the method further includes performing a fine-tuning process on the trained second model to obtain a fine-tuned second model; performing a compression and combination process on the fine-tuned second model to obtain a final second model; and using the final second model to predict utterance data.

[0106] Furthermore, the electronic device shown in FIG. 4 further includes a bus 102 and a communication interface 103, and the processor 101, the communication interface 103 and the memory 100 are connected via the bus 102.

[0107] Here, the memory 100 may include a high-speed random access memory (RAM) and may further include a non-volatile memory, such as at least one magnetic disk memory. At least one communication interface 103 (which may be wired or wireless) realizes a communication connection between the system network element and at least one other network element, and may use the Internet, a wide area network, a local network, a metro network, etc. The bus 102 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one double arrow is shown in FIG. 4, but this does not represent only one bus or only one type of bus.

[0108] The processor 101 may be an integrated circuit chip having signal processing capabilities. In the implementation process, each step of the above method can be completed by a hardware integrated logic circuit in the processor 101 or by instructions in software form. The processor 101 may be a general-purpose processor such as a central processing unit (CPU) or a network processor (NP), or may be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. Each method, step, and logic block diagram disclosed in the embodiments of the present invention can be realized or executed. The general-purpose processor may be a microprocessor, or any conventional processor, etc. The steps of the method disclosed in the embodiments of the present invention may be directly executed by a hardware decoding processor, or may be executed by a combination of hardware and software modules in the decoding processor. The software modules may be installed in a storage medium that is mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is installed in the memory 100, and the processor 101 reads information in the memory 100 and completes the steps of the method in the above embodiment in combination with its hardware.

[0109] An embodiment of the present invention further provides a computer-readable storage medium having computer-executable instructions stored thereon, which, when called and executed by a processor, causes the processor to implement the above-mentioned information identification method. For specific implementation, please refer to the embodiment of the method, and the description will be omitted here.

[0110] Specifically, the information identification method includes: acquiring first utterance data corresponding to a first game account in a target game; performing an abnormal content identification process on the first utterance data; and obtaining an identification result. If the identification result indicates that the first utterance data contains abnormal content, storing the first utterance data in a predetermined database, acquiring account information of the first game account, and storing the account information of the first game account in association with the first utterance data in the predetermined database. Inputting the account information of the first game account in the predetermined database into a pre-trained first model and obtaining a first output result, wherein the first output result is used to indicate the probability that the first game account is an abnormal account. If the probability corresponding to the first output result is greater than a predetermined probability threshold, inputting the first utterance data in the predetermined database into a pre-trained second model and obtaining a second output result, wherein the second output result includes at least one of information indicating that the utterance is normal, abnormal information corresponding to the first utterance data, contact information, and user identity information.

[0111] The above information identification method uses the account information and comment data of a game account to perform risk assessment and abnormality identification for the game account, thereby providing a safe gaming environment for users and helping to improve the user's gaming experience.

[0112] In an optional embodiment, performing an abnormal content identification process on the first utterance data and obtaining an identification result includes at least one of: matching the first utterance data with a predetermined regular expression library, obtaining a first matching result, and determining the first matching result as an identification result corresponding to the first utterance data, wherein if the first utterance data is matched with a regular expression in the regular expression library, the first matching result indicates that the first utterance data contains abnormal content; and matching the first utterance data with a predetermined keyword library, obtaining a second matching result, and determining the second matching result as a recognition result corresponding to the first utterance data, wherein if the first utterance data hits a keyword in the keyword library, the second matching result indicates that the first utterance data contains abnormal content.

[0113] In an alternative embodiment, the method is applied to a server, and the method further includes matching the first utterance data with a predetermined regular expression library by a regular matching module located in a game client corresponding to the target game, obtaining a first matching result, and sending the first matching result to the server.

[0114] In an alternative embodiment, the first model is obtained by training as follows: obtaining a first training set, the first training set including a plurality of first training samples, the first training samples including first preset account information and a first label corresponding to the first preset account information, the first label being used to indicate that the first preset account information is abnormal information; training an initial model based on the first training set to obtain an intermediate model; obtaining a second training set, the second training set including a plurality of second training samples; The second training sample includes a sample, and the second training sample includes second preset account information. The second training sample in the second training set is input into an intermediate model to obtain a model output result. A second label is set for the second training sample based on the model output result, and the second label is used to indicate a probability that the second preset account information in the second training sample is anomalous information. The intermediate model is trained based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set to obtain a first model.

[0115] In an optional embodiment, training an intermediate model based on the first training set, the second training set, and a second label corresponding to the second training sample in the second training set to obtain the first model includes: selecting a second training sample from the second training set whose second label is greater than a predetermined value, and setting the selected second training sample and the second label corresponding to the selected second training sample as a third training sample; and training an intermediate model based on the third training set and the first training set to obtain the first model.

[0116] In an alternative embodiment, the initial model includes a plurality of sub-models, and training the initial model based on the first training set to obtain intermediate models includes training each sub-model based on the first training set to obtain a plurality of intermediate sub-models, inputting second training samples in a second training set into the intermediate models and obtaining model output results, and setting second labels for the second training samples based on the model output results includes inputting the second training samples in the second training set into each intermediate sub-model, obtaining output results for each intermediate sub-model, and setting second labels for each intermediate sub-model. determining a second label of the second training sample based on an output result of the neural network model; training an intermediate model based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set to obtain the first model; training each intermediate sub-model based on the first training set, the second training set, and the second label corresponding to the second training sample in the second training set to obtain a plurality of trained intermediate sub-models; and determining the trained plurality of intermediate sub-models as the first model.

[0117] In an optional embodiment, determining a second label of the second training sample based on the output results of each intermediate sub-model includes either determining an average value of the output results of each intermediate sub-model as the second label of the second training sample, or determining a second label of the second training sample based on the output results of each intermediate sub-model and output weights corresponding to each intermediate sub-model, wherein the output weights corresponding to the intermediate sub-models are determined based on second predetermined account information in the second training sample.

[0118] In an alternative embodiment, the first model includes a plurality of trained intermediate sub-models, and inputting the account information of the first game account in the preset database into the pre-trained first model and obtaining a first output result includes inputting the account information of the first game account in the preset database into a plurality of trained intermediate sub-models, respectively, and obtaining an output result corresponding to each trained intermediate sub-model; and determining an average value of the output results corresponding to each trained intermediate sub-model as the first output result, or determining the first output result based on the output results corresponding to each trained intermediate sub-model and output weights corresponding to each trained intermediate sub-model, wherein the output weights corresponding to the trained intermediate sub-models are determined based on information included in the account information of the first game account.

[0119] In an optional embodiment, the second model is obtained by training as follows: a fourth training set is obtained, the fourth training set includes a plurality of fourth training samples, the fourth training samples include predetermined utterance data and a fourth label corresponding to the predetermined utterance data, the fourth label is used to indicate whether the predetermined utterance data is an abnormal utterance, and if the predetermined utterance is an abnormal utterance, the fourth label includes at least one of anomaly information, contact information, and user identity information corresponding to the predetermined utterance data; and the second model is trained based on the fourth training set and a predetermined dynamic learning rate scheduling policy, and a trained second model is obtained.

[0120] In an optional embodiment, the preset utterance data included in the fourth training sample is obtained by dividing the original utterance data using a sliding window, and when dividing the original utterance data, there is an overlapping portion between the divided utterance data in the front and rear stages.

[0121] In an alternative embodiment, the method further includes performing a fine-tuning process on the trained second model to obtain a fine-tuned second model; performing a compression and combination process on the fine-tuned second model to obtain a final second model; and using the final second model to predict utterance data.

[0122] The functions may be implemented as software functional units and stored in a computer-readable storage medium when sold or used as a standalone product. Based on this understanding, it is understood that the technical solution of the present invention, or a part of the technical solution that essentially contributes to the prior art, may be embodied in the form of a software product stored in a storage medium containing a number of instructions that enable a computer device (which may be a personal computer, a terminal device, a network device, etc.) to execute all or part of the methods described in various embodiments of the present invention. The storage medium may be various media capable of storing program code, such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a disk, or a CD-ROM.

[0123] In describing the present invention, orientations or positional relationships indicated by terms such as "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer" are based on orientations or positional relationships shown in the drawings and are intended merely to facilitate and simplify the description of the present invention, and do not indicate or imply that the devices or components referred to must have a particular orientation or be configured and operated in a particular orientation, and therefore cannot be understood as limiting the present invention. Furthermore, the terms "first," "second," and "third" are merely for illustrative purposes and cannot be understood as indicating or implying relative importance.

[0124] Finally, it should be explained that the above examples are only specific embodiments of the present invention, and are intended to illustrate the technical solutions of the present invention, but are not limited thereto, and the protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above examples, those skilled in the art who are familiar with the technical field can still modify or easily change the technical solutions described in the above examples within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features therein, and these modifications, changes or substitutions should be included in the protection scope of the present invention without departing from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention. Therefore, the protection scope of the present invention is subject to the protection scope of the claims.

Claims

1. acquiring first utterance data corresponding to a first game account in the target game, performing an abnormal content identification process on the first utterance data, and acquiring an identification result; If the identification result indicates that the first utterance data contains abnormal content, storing the first utterance data in a preset database, obtaining account information of the first game account, and storing the account information of the first game account in association with the first utterance data in the preset database; inputting account information of the first game account in the preset database into a pre-trained first model to obtain a first output result, wherein the first output result is used to indicate a probability that the first game account is an anomalous account; If the probability corresponding to the first output result is greater than a predetermined probability threshold, inputting the first utterance data in the predetermined database into a pre-trained second model to obtain a second output result, wherein the second output result includes at least one of information indicating that the utterance is normal, abnormal information corresponding to the first utterance data, contact information, and user identity information.

1. An information identification method comprising:

2. performing an abnormal content identification process on the first utterance data and obtaining an identification result, Matching the first utterance data with a predetermined regular expression library, obtaining a first matching result, and determining the first matching result as an identification result corresponding to the first utterance data, wherein when the first utterance data is matched with a regular expression in the regular expression library, the first matching result indicates that the first utterance data contains abnormal content; The method includes at least one of: matching the first utterance data with a preset keyword library; obtaining a second matching result; and determining the second matching result as a recognition result corresponding to the first utterance data, wherein if the first utterance data matches a keyword in the keyword library, the second matching result indicates that the first utterance data contains abnormal content.

2. The information identification method according to claim 1.

3. An information identification method applied to a server, comprising: and further including: matching the first utterance data with a predetermined regular expression library by a regular matching module disposed in a game client corresponding to the target game, obtaining a first matching result, and transmitting the first matching result to the server.

3. The information identification method according to claim 2.

4. The first model is obtained by training it as follows: A first training set is obtained, the first training set includes a plurality of first training samples, the first training samples include first preset account information and a first label corresponding to the first preset account information, and the first label is used to indicate that the first preset account information is anomalous information; training an initial model based on the first training set to obtain an intermediate model; obtaining a second training set, the second training set including a plurality of second training samples, the second training samples including second preset account information; inputting a second training sample in the second training set into the intermediate model to obtain a model output result; and setting a second label for the second training sample based on the model output result, wherein the second label is used to indicate a probability that second predetermined account information in the second training sample is anomalous information; training the intermediate model based on the first training set, the second training set, and a second label corresponding to a second training sample in the second training set to obtain the first model; 2. The information identification method according to claim 1.

5. Training the intermediate model based on the first training set, the second training set, and second labels corresponding to second training samples in the second training set to obtain the first model includes: selecting, from the second training set, a second training sample whose second label is greater than a predetermined value, and setting the selected second training sample and a second label corresponding to the selected second training sample as a third training sample; training the intermediate model based on a third training set and the first training set to obtain the first model.

5. The information identification method according to claim 4.

6. the initial model includes a plurality of sub-models; training an initial model based on the first training set to obtain intermediate models includes training each of the sub-models based on the first training set to obtain a plurality of intermediate sub-models; inputting second training samples in the second training set into the intermediate model, obtaining model output results, and setting second labels for the second training samples based on the model output results includes inputting second training samples in the second training set into each of the intermediate sub-models, obtaining output results for each of the intermediate sub-models, and determining second labels for the second training samples based on the output results for each of the intermediate sub-models; Training the intermediate model based on the first training set, the second training set, and second labels corresponding to second training samples in the second training set to obtain the first model includes training each of the intermediate sub-models based on the first training set, the second training set, and second labels corresponding to second training samples in the second training set to obtain a plurality of trained intermediate sub-models, and determining the trained plurality of intermediate sub-models as the first model.

6. The information identification method according to claim 5.

7. determining a second label of the second training sample based on an output result of each of the intermediate sub-models; determining an average value of the output results of each of the intermediate sub-models as a second label of the second training sample; determining a second label of the second training sample based on an output result of each of the intermediate sub-models and an output weight corresponding to each of the intermediate sub-models, wherein the output weight corresponding to the intermediate sub-model is determined based on second predetermined account information of the second training sample.

7. The information identification method according to claim 6.

8. the first model includes a plurality of trained intermediate sub-models; inputting account information of the first game account in the preset database into a pre-trained first model and obtaining a first output result; inputting account information of the first game account in the preset database into the trained intermediate sub-models respectively, and obtaining output results corresponding to each trained intermediate sub-model; determining an average value of output results corresponding to each trained intermediate sub-model as the first output result, or determining the first output result based on the output results corresponding to each trained intermediate sub-model and an output weight corresponding to each trained intermediate sub-model, wherein the output weight corresponding to the trained intermediate sub-model is determined based on information included in account information of the first game account.

2. The information identification method according to claim 1.

9. The second model is obtained by training it as follows: a fourth training set is obtained, the fourth training set includes a plurality of fourth training samples, the fourth training samples include predetermined utterance data and a fourth label corresponding to the predetermined utterance data, the fourth label is used to indicate whether the predetermined utterance data is an abnormal utterance, and if the predetermined utterance is an abnormal utterance, the fourth label includes at least one information among abnormal information, contact information, and user identity information corresponding to the predetermined utterance data; Training the second model based on the fourth training set and a preset dynamic learning rate scheduling policy to obtain a trained second model.

2. The information identification method according to claim 1.

10. The predetermined utterance data included in the fourth training sample is obtained by dividing the original utterance data using a sliding window, and when dividing the original utterance data, there is an overlapping portion between the divided utterance data in the front part and the divided utterance data in the rear part.

10. The information identification method according to claim 9.

11. performing a fine-tuning process on the trained second model to obtain a fine-tuned second model; performing a compression and combination process on the fine-tuned second model to obtain a final second model, and using the final second model to predict utterance data.

10. The information identification method according to claim 9.

12. An information identification device comprising: a data acquisition module, an information storage module, a first identification module, and a second identification module, the data acquisition module is configured to acquire first utterance data corresponding to a first game account in the target game, perform an abnormal content identification process on the first utterance data, and acquire an identification result; the information storage module is configured, when the identification result indicates that the first utterance data contains abnormal content, to store the first utterance data in a preset database, obtain account information of the first game account, and associate the account information of the first game account with the first utterance data and store it in the preset database; The first identification module is configured to input account information of the first game account in the preset database into a pre-trained first model to obtain a first output result, and the first output result is used to indicate a probability that the first game account is an abnormal account; The second identification module is configured to input the first utterance data in the preset database into a pre-trained second model to obtain a second output result when the probability corresponding to the first output result is greater than a preset probability threshold, and the second output result includes at least one of information that the utterance is normal, abnormal information corresponding to the first utterance data, contact information, and user identity information. An information identification device characterized by:

13. The information identification method according to any one of claims 1 to 11 includes a processor and a memory, and the memory stores machine-executable instructions that can be executed by the processor. The processor executes the machine-executable instructions to realize the information identification method according to any one of claims 1 to 11. An electronic device characterized by:

14. a computer-readable storage medium storing computer-executable instructions, which, when called and executed by a processor, cause the processor to implement the information identification method of any one of claims 1 to 11; A computer-readable storage medium comprising:

Citation Information

Patent Citations

  • Game data monitoring system and game data monitoring method

    CN113318454A

  • Method and device for identifying abnormal account in game, electronic equipment and storage medium

    CN113440856A

  • SYSTEM AND METHOD FOR VERIFYING GAME PLAY-RELATED ACTIVITY - Patent application

    JP2023545108A

  • Automatic classification and reporting of inappropriate language in online applications

    US20210370188A1

  • Offensive chat filtering using machine learning models

    US20220284884A1