Multi-dimension-based information identification method and apparatus, and electronic device

By conducting multi-dimensional analysis of the information exchange data between the risky sending and receiving ends, distinguishing between one-way and interactive information, and using a trained model to calculate probability values, the problem of accuracy and comprehensiveness in abnormal SMS identification in existing technologies has been solved, achieving more efficient identification of abnormal information.

CN120995299APending Publication Date: 2025-11-21CHINA MOBILE GRP HENAN CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510939643.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing abnormal SMS identification methods based on large models only judge the content of a single SMS message, resulting in low accuracy and comprehensiveness, and failing to provide a comprehensive understanding of the overall situation of the SMS message.

Method used

By conducting multi-dimensional analysis of the information exchange data between the risky sending end and multiple receiving ends, we can distinguish between one-way information and interactive information, and use a trained information recognition model to calculate the probability value of the information to identify abnormal information data.

Benefits of technology

It improves the accuracy and comprehensiveness of abnormal SMS identification, enabling it to identify individual messages as normal but determine them as abnormal after combining multiple messages, thus enhancing the ability to identify fraudulent activities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995299A_ABST
    Figure CN120995299A_ABST
Patent Text Reader

Abstract

The invention provides a multi-dimension-based information identification method and device and electronic equipment, and relates to the technical field of data processing, and the method comprises the steps: carrying out the data analysis processing of a plurality of pieces of information data transmitted by an identified risk transmitting end, determining a receiving end, and obtaining the information interaction data between the risk transmitting end and the receiving end; performing classification processing on the multiple pieces of information data according to information interaction data to obtain one-way information data and interaction information data; inputting the one-way information data, and / or the interaction information data and the interaction information into a trained information identification model for information identification processing to obtain an information probability value; and determining the information data with the information probability value greater than a preset probability threshold as abnormal information data. The condition that a single piece of information is normal but can be determined as abnormal information after multiple pieces of information are integrated can be effectively identified, so that the accuracy and comprehensiveness of information identification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of data processing, and particularly relates to a method and device for information recognition based on multiple dimensions, and an electronic device. BACKGROUND

[0002] With the rapid development of mobile communication technology and the popularity of smart phones, the security problem of short message (Short Message Service, SMS) as one of the basic communication methods is increasingly prominent. Short message may become a major threat to property and information due to its wide range of dissemination, low cost and varied forms. Therefore, it is crucial to effectively identify and intercept abnormal short messages to protect the security of communication networks and the rights and interests of users.

[0003] Due to the wide application of large language models (Large Language Models, LLMs), the method based on large models for abnormal short message recognition can exhibit good results in identifying new and variant abnormal short messages containing inducible language, false information or high-risk links. However, the existing method for intelligent identification of abnormal short messages based on large models only judges single short message content, and cannot fully understand the overall situation of the short message, resulting in low accuracy and comprehensiveness in identifying abnormal short messages.

[0004] Therefore, how to improve the accuracy and comprehensiveness of abnormal short message recognition is a problem to be solved at present. SUMMARY

[0005] The present disclosure provides a method and device for information recognition based on multiple dimensions, and an electronic device. The main purpose is to solve the problem of how to improve the accuracy and comprehensiveness of abnormal short message recognition.

[0006] According to a first aspect of the present disclosure, a method for information recognition based on multiple dimensions is provided, which comprises:

[0007] performing data analysis and processing on the identified multiple information data sent by the risk sending end, determining the respective receiving ends corresponding to the multiple information data, and obtaining information interaction data between the risk sending end and the respective receiving ends corresponding to the multiple information data;

[0008] performing classification processing on the multiple information data according to the information interaction data, to obtain unidirectional class information data and interactive class information data, wherein the unidirectional class information data is information data in the multiple information data without reply information, and the interactive class information data is information data in the multiple information data with reply information;

[0009] inputting the unidirectional information data and / or the interactive information data and the interaction information corresponding to the interactive information data into the trained information recognition model to perform information recognition processing, to obtain information probability values corresponding to the plurality of information data respectively, wherein the trained information recognition model comprises at least recognition parameters of the unidirectional information data and recognition parameters of the interactive information data;

[0010] information data with an information probability value greater than a preset probability threshold is determined as abnormal information data.

[0011] Optionally, the information interaction data between the risk information sender and the plurality of information data corresponding receiving ends comprises:

[0012] determining a target receiving end that sends reply information to the risk information sender from the plurality of information data corresponding receiving ends;

[0013] obtaining information interaction data between the target receiving end and the risk information sender.

[0014] Optionally, the classification processing of the plurality of information data according to the information interaction data comprises:

[0015] determining information data with reply information from the plurality of information data according to the information interaction data;

[0016] determining the information data with the reply information as the interactive information data, and obtaining interaction information corresponding to the interactive information data from the information interaction data;

[0017] determining information data without the reply information as the unidirectional information data.

[0018] Optionally, before the data analysis processing of the plurality of information data sent by the identified risk information sender and the determination of the plurality of information data corresponding receiving ends, the method further comprises:

[0019] obtaining training unidirectional information data, training interactive information data and training interaction information corresponding to the training interactive information data;

[0020] setting a label length parameter of the to-be-trained model as a first value to obtain a first set model, inputting the training unidirectional information data into the first set model, and performing model training processing on the first set model based on a preset loss function and a preset training algorithm to obtain a first information recognition model;

[0021] Set the mark length parameter of the first information recognition model to a second value to obtain a second set model, input the training interactive information and the training interactive information corresponding to the training interactive information into the second set model, and perform model training processing on the second set model based on a preset loss function and a preset training algorithm to obtain a trained information recognition model.

[0022] Optionally, after inputting the one-way information data, and / or the interactive information and the interactive information corresponding to the interactive information into the trained information recognition model for information recognition processing to obtain information probability values corresponding to the plurality of information data respectively, the method further comprises:

[0023] Determining whether the information probability value is greater than a preset probability threshold;

[0024] In a case where it is determined that the information probability value is not greater than the preset probability threshold, determining the information data whose information probability value is not greater than the preset probability threshold as normal information data;

[0025] In a case where it is determined that the information probability value is greater than the preset probability threshold, determining the information data whose information probability value is greater than the preset probability threshold as abnormal information data.

[0026] Optionally, before performing data analysis processing on the plurality of information data sent by the identified risk sender to determine the plurality of information data corresponding to the receiving end respectively, the method further comprises:

[0027] Performing value assignment processing on the plurality of sending ends to obtain initial feature values corresponding to the plurality of sending ends respectively;

[0028] Obtaining sender attribute data corresponding to the plurality of sending ends respectively, and adjusting the initial feature values according to the sender attribute data and a preset risk condition to obtain target feature values corresponding to the plurality of sending ends respectively, and determining a maximum feature value in the plurality of target feature values;

[0029] Performing data normalization processing on the plurality of target feature values and the maximum feature value respectively to obtain first risk scores corresponding to the plurality of sending ends respectively;

[0030] Determining the sending end whose first risk score is greater than a first preset risk threshold as a risk sender.

[0031] Optionally, after inputting the one-way information data, and / or the interactive information and the interactive information corresponding to the interactive information into the trained information recognition model for information recognition processing to obtain information probability values corresponding to the plurality of information data respectively, the method further comprises:

[0032] acquire a first information quantity corresponding to a plurality of information data, a second information quantity of cross-region information in the plurality of information data, and a third information quantity of reply information received by the risk signaling end;

[0033] perform calculation processing according to the first information quantity and the second information quantity to obtain a cross-region information proportion, and perform calculation processing according to the first information quantity and the third information quantity to obtain an information sending-receiving ratio;

[0034] perform data calculation processing according to the first information quantity, the cross-region information proportion, and the information sending-receiving ratio to obtain a second risk score corresponding to the risk signaling end;

[0035] perform calculation processing according to the first risk score, a preset weight, and the second risk score to obtain a comprehensive risk score corresponding to the risk signaling end, and perform shutdown processing on the risk signaling end in a case where the comprehensive risk score is greater than a second preset risk threshold.

[0036] Optionally, the performing calculation processing according to the first risk score, the preset weight, and the second risk score to obtain the comprehensive risk score corresponding to the risk signaling end comprises:

[0037] The calculation of the comprehensive risk score adopts the following formula:

[0038]

[0039] wherein, X is the comprehensive risk score, M i is an information probability value corresponding to the i-th information data in the plurality of information data, N is the first information quantity, W is the cross-region information proportion, and F is the information sending-receiving ratio.

[0040] According to a second aspect of the present disclosure, an apparatus based on multi-dimensional information recognition is provided, comprising:

[0041] an analysis unit configured to perform data analysis processing on a plurality of information data sent by a recognized risk signaling end, and determine a respective corresponding receiving end of the plurality of information data;

[0042] an acquisition unit configured to acquire information interaction data between the risk signaling end and the respective corresponding receiving end of the plurality of information data;

[0043] a classification unit configured to perform classification processing on the plurality of information data according to the information interaction data to obtain unidirectional class information data and interactive class information data, wherein the unidirectional class information data is information data in the plurality of information data that does not have reply information, and the interactive class information data is information data in the plurality of information data that has reply information;

[0044] The recognition unit is configured to input the unidirectional information data, the interaction information data, and interaction information corresponding to the interaction information data into the trained information recognition model to perform information recognition processing, and obtain information probability values corresponding to the information data respectively.

[0045] The determination unit is configured to determine information data with an information probability value greater than a preset probability threshold as abnormal information data.

[0046] Optionally, the acquisition unit is further configured to:

[0047] determine a target receiving end from receiving ends corresponding to the information data respectively, which sends reply information to the risk sending end;

[0048] acquire information interaction data between the target receiving end and the risk sending end.

[0049] Optionally, the classification unit is further configured to:

[0050] determine information data with reply information from the information data according to the information interaction data;

[0051] determine the information data with the reply information as interaction information data, and acquire interaction information corresponding to the interaction information data from the information interaction data;

[0052] determine information data without the reply information as unidirectional information data.

[0053] Optionally, the acquisition unit is further configured to acquire training unidirectional information data, training interaction information data, and training interaction information corresponding to the training interaction information data.

[0054] The device for multi-dimensional information recognition further comprises:

[0055] The training unit is configured to set a mark length parameter of a to-be-trained model as a first value to obtain a first set model, input the training unidirectional information data into the first set model, perform model training processing on the first set model based on a preset loss function and a preset training algorithm, and obtain a first information recognition model.

[0056] The training unit is further configured to set the mark length parameter of the first information recognition model as a second value to obtain a second set model, input the training interaction information data and the training interaction information corresponding to the training interaction information data into the second set model, perform model training processing on the second set model based on the preset loss function and the preset training algorithm, and obtain the trained information recognition model.

[0057] Optionally, the determining unit is further configured to:

[0058] determine whether the information probability value is greater than a preset probability threshold value;

[0059] in a case where it is determined that the information probability value is not greater than the preset probability threshold value, determine information data for which the information probability value is not greater than the preset probability threshold value as normal information data;

[0060] in a case where it is determined that the information probability value is greater than the preset probability threshold value, determine information data for which the information probability value is greater than the preset probability threshold value as abnormal information data.

[0061] Optionally, the device for multi-dimensional information recognition further comprises:

[0062] a processing unit configured to perform value assignment processing on the multiple signaling ends to obtain initial feature values corresponding to the multiple signaling ends respectively;

[0063] the processing unit is further configured to obtain signaling end attribute data corresponding to the multiple signaling ends respectively, and perform adjustment processing on the initial feature values according to the signaling end attribute data and a preset risk condition to obtain target feature values corresponding to the multiple signaling ends respectively, and determine a maximum feature value in the multiple target feature values;

[0064] the processing unit is further configured to perform data normalization processing on the multiple target feature values and the maximum feature value respectively to obtain first risk scores corresponding to the multiple signaling ends respectively;

[0065] the processing unit is further configured to determine a signaling end with a first risk score greater than a first preset risk threshold value as a risk signaling end.

[0066] Optionally, the obtaining unit is further configured to obtain a first information quantity corresponding to the multiple information data, a second information quantity of cross-region information in the multiple information data, and a third information quantity of reply information received by the risk signaling end;

[0067] the processing unit is further configured to perform calculation processing according to the first information quantity and the second information quantity to obtain a cross-region information proportion, and perform calculation processing according to the first information quantity and the third information quantity to obtain an information transceiving ratio;

[0068] the processing unit is further configured to perform data calculation processing according to the first information quantity, the cross-region information proportion, and the information transceiving ratio to obtain a second risk score corresponding to the risk signaling end;

[0069] the processing unit is further configured to perform calculation processing according to the first risk score, a preset weight, and the second risk score to obtain a comprehensive risk score corresponding to the risk signaling end, and perform shutdown processing on the risk signaling end in a case where the comprehensive risk score is greater than a second preset risk threshold value.

[0070] Optionally, the processing unit is further configured to:

[0071] The comprehensive risk score is calculated using the following formula:

[0072]

[0073] wherein X is the comprehensive risk score, M i is the information probability value corresponding to the i-th information data in the plurality of information data, N is the first information quantity, W is the cross-regional information proportion, and F is the information sending and receiving ratio.

[0074] According to a third aspect of the present disclosure, an electronic device is provided, comprising:

[0075] at least one processor; and

[0076] a memory in communication with the at least one processor; wherein

[0077] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of information identification based on multiple dimensions according to the first aspect.

[0078] According to a fourth aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to perform the method of information identification based on multiple dimensions according to the first aspect.

[0079] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method of information identification based on multiple dimensions according to the first aspect.

[0080] The method and device of information identification based on multiple dimensions provided by the present disclosure, by analyzing all information or multiple information sent by the sending terminal based on the behavior mode of the sending terminal, determining the category of the information, i.e. the single-way category information and the interactive category information, directly identifying the information itself for the single-way category information, identifying the information itself and the corresponding interactive information for the interactive category information, distinguishing the single-way category information and the interactive category information, identifying the information content, and combining the information itself and the interactive information replied by the receiving terminal, the accuracy of information identification is improved, and the case that a single information is normal but multiple information can be determined as abnormal information can be effectively identified, thereby improving the accuracy and comprehensiveness of information identification.

[0081] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0082] The accompanying drawings are used to better understand the present scheme, and do not constitute a limitation on the present disclosure. Among them:

[0083] Figure 1 A flowchart of a method for information recognition based on multiple dimensions provided by an embodiment of the present disclosure;

[0084] Figure 2 A structural diagram of an apparatus for information recognition based on multiple dimensions provided by an embodiment of the present disclosure;

[0085] Figure 3 A structural diagram of another apparatus for information recognition based on multiple dimensions provided by an embodiment of the present disclosure;

[0086] Figure 4 A schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0087] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in order to be clear and concise, descriptions of well-known functions and structures are omitted in the following description.

[0088] Embodiments of the present disclosure relate to a method for information recognition based on multiple dimensions, which realizes accurate identification of abnormal information by comprehensively analyzing interactive behavior characteristics between a risk signaling end and multiple receiving ends. The method first performs in-depth correlation analysis on original information data sent by the risk signaling end, then classifies information types according to interactive mode differences, and finally completes risk determination through identification models adapted to different interactive characteristics.

[0089] The method and apparatus for information recognition based on multiple dimensions, and the electronic device of an embodiment of the present disclosure are described below with reference to the accompanying drawings.

[0090] Figure 1 A flowchart of a method for information recognition based on multiple dimensions provided by an embodiment of the present disclosure.

[0091] As Figure 1 shown, the method includes the following steps:

[0092] Step 101, performing data analysis and processing on multiple information data sent by the identified risk signaling end, determining respective receiving ends corresponding to the multiple information data, and obtaining information interaction data between the risk signaling end and the respective receiving ends corresponding to the multiple information data.

[0093] In the embodiments of the present application, the risk signaling end refers to an information sending subject marked as suspicious after preliminary risk assessment, and the risk attribute of the risk signaling end can be determined through dimensions such as historical violation records, equipment feature abnormalities, or behavior pattern deviations.

[0094] Among them, the multiple information data sent by the risk signaling end need to be structured, and the information data includes but is not limited to text content, sending timestamp, transmission protocol type, and other original information metadata. The mapping relationship between the information data and the receiving end is established through data analysis and processing.

[0095] The receiving end is an information receiving party entity, and the identity of the receiving end can be a phone number, a user account, or a device code. On this basis, the information interaction data is dynamically obtained, and the information interaction data is a two-way communication record set formed between the risk signaling end and each receiving end, which not only includes the initial sent information data, but also covers the reply content and the corresponding time sequence relationship generated by the receiving end.

[0096] In step 102, the multiple information data are classified according to the information interaction data, and unidirectional class information data and interactive class information data are obtained, wherein the unidirectional class information data is information data in which there is no reply information in the multiple information data, and the interactive class information data is information data in which there is reply information in the multiple information data.

[0097] In the embodiments of the present application, the multiple information data are divided into two types of communication modes with significant behavior differences according to whether there is a reply behavior in the information interaction data.

[0098] The unidirectional class information data refers to an information instance that does not trigger any reply of the receiving end after being sent by the risk signaling end. The unidirectional class information data usually represents a non-interactive communication scenario such as unidirectional notification, advertisement push, or automatic alarm. The corresponding interactive class information data refers to information data that triggers active reply or forms multiple rounds of dialogue of the receiving end, and the feature is that the communication parties establish an information exchange channel.

[0099] In the classification process, by analyzing the response delay, dialogue round, and content correlation degree in the information interaction data, it is automatically identified whether there is an effective reply behavior, so as to realize accurate separation of the two types of information data.

[0100] Specifically, the method for information classification is, for example:

[0101] A certain signaling person (risk signaling end) and each of his receiving persons (receiving end) are taken as a group respectively. The signaling person is A, and the receiving persons are B1, B2, B3, …, BN. A sends a message to the receiving persons B1, B2, B3, …, BN.

[0102] If B does not reply to A, the short message is defined as single-way type information data.

[0103] If B1, B2, B3, BN reply to A, and A has the interaction of short messages, the short message is defined as interactive type information data. Then AB1, AB2, AB3, …, ABN are paired respectively, and the content of each pair of interactive short messages is judged.

[0104] For A sends multiple short messages to B, but B does not reply. This case can be identified as single-way type or interactive type, and at the same time, an information data with the same content can be interactive type information or single-way type information in different receiving ends corresponding to the risk sending end. For example, in the AB1 combination, B1 does not reply to information i, so information i in the AB1 combination is single-way type information data. In the AB2 combination, B2 replies to information i, so information i in the AB2 combination is interactive type information data.

[0105] In step 103, the single-way type information data and / or the interactive type information data and the interactive information corresponding to the interactive type information data are input into the trained information recognition model for information recognition processing to obtain information probability values corresponding to the information data respectively.

[0106] In the embodiments of the present disclosure, the trained information recognition model is used to analyze the classified information data in parallel. The trained information recognition model can include two independent functional modules: the first module is used to process single-way type information data, and the recognition parameters thereof are optimized for the semantic features of a single information; the second module is used to analyze interactive type information data and the interactive information corresponding thereto, and the recognition parameters thereof focus on learning the context association features of the dialogue.

[0107] The information probability value is a core risk indicator output by the model, and the numerical interval thereof is [0, 1], which represents the possibility degree of a single information being judged as abnormal. It needs to be particularly pointed out that the model can be flexibly deployed according to actual scene requirements: the first module is enabled when only single-way type information data exists, the second module is enabled when interactive type information data exists, and the dual-module collaborative analysis mechanism is started when the two types of data coexist.

[0108] In step 104, the information data with an information probability value greater than a preset probability threshold is determined as abnormal information data.

[0109] In the embodiments of the present disclosure, a dynamic judgment boundary is established by setting a preset probability threshold. The preset probability threshold is a threshold set by the user, for example, 0.5, etc. The preset probability threshold can be configured according to the security level requirements of the actual business scene.

[0110] The information probability value output by the model is compared with a preset probability threshold in real time, and when the information probability value of certain information data exceeds the preset probability threshold, the information data is marked as abnormal information data. The information recognition method of the present disclosure effectively solves the misjudgment problem caused by ignoring the interactive behavior characteristics, and has a significant recognition effect on the hidden fraudulent behavior gradually induced through multiple rounds of dialogue in the interactive information.

[0111] By fusing communication behavior pattern analysis and content semantic recognition, a multi-dimensional information risk assessment system is constructed. Compared with a single content analysis method, the following technical effects are achieved: through interactive mode classification, the recognition model can adapt to the behavior characteristics of different fraudulent means; through dialogue context analysis of interactive information, the detection ability for gradual abnormalities is significantly improved; and through a dynamic judgment mechanism based on a probability threshold, the adaptability of the system to new fraudulent strategies is enhanced. The coverage and judgment accuracy of abnormal information recognition are improved, and complex fraudulent behaviors that evade traditional detection means are effectively suppressed.

[0112] The multi-dimensional information recognition method provided by the present disclosure comprehensively analyzes all information or multiple information sent by the information sender based on the behavior pattern of the information sender, determines the category of the information, i.e., single-way information and interactive information, directly recognizes the information itself for single-way information, recognizes the information itself and the corresponding interactive information for interactive information, and improves the accuracy of information recognition by distinguishing single-way information and interactive information, recognizing the information content, and combining the information itself and the interactive information replied by the information receiver. The method can effectively identify the case that a single information is normal but multiple information can be determined as abnormal information, thereby improving the accuracy and comprehensiveness of information recognition.

[0113] In an implementation manner of an embodiment of the present disclosure, when the information interaction data is acquired, the following implementation manners can be used, but are not limited to: determining a target information receiver that sends reply information to the risk information sender among the information receivers corresponding to the multiple information data; and acquiring the information interaction data between the target information receiver and the risk information sender.

[0114] In an embodiment of the present disclosure, the dynamic acquisition of information interaction data is achieved by establishing a behavior response graph between the risk information sender and the information receiver, accurately capturing the interactive communication scene, identifying a specific information receiver with active feedback behavior, and extracting a complete two-way dialogue chain accordingly.

[0115] The target receiving end is a subset of receiving ends with reply information behavior, which sends reply information to the risk sending end. The information flow of the reply information is opposite to the original sending direction of the information data (i.e. initiated by the receiving end to the risk sending end); the content of the reply information has semantic relevance with the original information data (such as: reply to inquiry, feedback to notification, etc.).

[0116] Regarding the determination of the reply information, the identification can be completed through the time stamp reverse tracking mechanism of the communication log: first, lock the information data sent by the risk sending end and its corresponding receiving end identifier, and then scan the reverse information flow sent by the receiving end within a preset time window (such as: 24-72 hours). When it is detected that the receiving address matches the risk sending end identifier and the content contains the conversation continuation feature (such as: quoting the original text, using the same conversation), it is confirmed that this receiving end becomes the target receiving end.

[0117] The information interaction data is the complete conversation information formed between the risk sending end and a single target receiving end, and the information interaction data at least includes: the initial information data of the risk sending end, the reply information of the target receiving end, and all associated information subsequently generated by both parties. The extraction process can use but is not limited to the conversation tree construction algorithm, taking the information data first sent by the risk sending end as the root node, taking the reply of the target receiving end as the first level child node, and the subsequently alternately generated information forming multiple levels of branches in chronological order. At the same time, the conversation segmentation processing can be performed on the non-continuous conversation scene (such as: the conversation initiated again after several days interval), to ensure that each piece of information interaction data only contains a logically coherent conversation sequence. In the finally generated structured data packet, each piece of information is attached with a time stamp accurate to milliseconds, a sender / receiver identity tag and an original content code.

[0118] Since the fraudulent behavior often has the timeliness feature, when the information interaction data is acquired, an incremental acquisition strategy can also be used, when a target receiving end is newly identified, a backtracking scan will be automatically triggered, not only to acquire the information interaction data that already exists at the current time, but also to continuously monitor the subsequent dynamics of the conversation channel. Any newly added reply information will be appended to the corresponding information interaction data set in real time. Through the dynamic updating mechanism, even if the risk sending end adopts the long-term criminal mode of phishing-waiting-secondary induction, the complete interaction behavior chain can still be captured.

[0119] By accurately positioning the active responder and extracting structured conversation data, high-quality input sources are provided for subsequent classification and identification. Further, by focusing on the target recipient, the data processing amount is greatly reduced, and the invalid resource consumption of non-responding recipients is avoided. By conversation tree construction algorithm, the context logical relationship of interactive behavior is completely retained, and the semantic break problem caused by time sequence disorder is overcome. By incremental acquisition strategy, the delay response characteristics of fraudulent behavior are effectively dealt with, and the omission of key evidence is prevented. The accuracy and timeliness of information interaction data acquisition are improved, and a solid foundation is laid for subsequent risk judgment.

[0120] In an implementation manner of the embodiment of the present disclosure, when the plurality of information data is classified and processed, the following manner can be used, but is not limited to: determining information data with reply information in the plurality of information data according to the information interaction data; determining the information data with the reply information as interactive information data, and obtaining interactive information corresponding to the interactive information data from the information interaction data; and determining information data without reply information as unidirectional information data.

[0121] In the embodiment of the present disclosure, the conversation graph in the information interaction data can be used to verify whether there is reply information. When the information sent by a certain recipient and the information data previously sent by the risk sender form a reply relationship, and the two are bound in the same communication thread through the conversation identifier (Identifier, ID) or time sequence, it is confirmed that the information data has valid reply information. Verification through the conversation graph can effectively exclude the interference of irrelevant information, for example, a new conversation initiated by the recipient will not be mistaken as a reply behavior.

[0122] When it is confirmed that a certain information data has reply information, it is dynamically labeled as interactive information data. This classification label not only identifies that the information data triggers a two-way conversation, but also triggers the deep extraction of associated data. Then, complete interactive information can be obtained from the information interaction data. The interactive information includes multiple levels of conversation units: original trigger information (the content first sent by the risk sender), first reply information (the first feedback of the target recipient), and all subsequent associated conversation content.

[0123] In the extraction process of the interactive information, the conversation thread reconstruction technology can be used to recombine the dispersedly stored information into a logically coherent conversation chain according to the time line through the session identifier, so as to ensure the integrity of the context semantics. For scenes containing multiple rounds of alternation, the initiator identifier and accurate timestamp of each round of conversation are recorded to form a tree-shaped conversation structure with time sequence markers.

[0124] For information data that does not trigger any valid reply, it is classified into the one-way type information data category. The communication life cycle of the one-way type information data ends with the first sending behavior, and the receiving end does not generate any semantically associated feedback. For the one-way type information data, a negative feedback mechanism can be used for judgment: when certain information data is not responded to by any target receiving end within a preset monitoring period, and no extended branch is formed in the dialogue graph, it is automatically marked as one-way type information data. At the same time, an anti-misjudgment strategy is set: for reply information that arrives after exceeding the monitoring period due to network delay and the like, the classification result is dynamically updated through an asynchronous processing mechanism to ensure the real-time accuracy of the classification state.

[0125] By accurately identifying the characteristics of information interaction behaviors, high-quality input data sources are provided for subsequent differentiated recognition models. By using double verification standards for reply information, the misclassification rate is significantly reduced, and accidental communication is avoided from being included in the interaction category; by structuring the extraction of interaction information, the dialogue context is completely preserved, providing a key basis for recognizing progressive fraud behaviors; and by using a dynamic collection mechanism, the purity of one-way type information is ensured, preventing invalid data from interfering with the recognition model. These improvements enable the classification result to more accurately reflect the essential characteristics of information dissemination, and establish a reliable data foundation for subsequent risk judgment.

[0126] In an implementation manner of the embodiment of the present disclosure, before data analysis is performed on the plurality of information data, the following methods can also be used, but are not limited to: obtaining training one-way type information data, training interaction type information data, and training interaction information corresponding to the training interaction type information data; setting a label length parameter of a to-be-trained model as a first value to obtain a first set model, inputting the training one-way type information data into the first set model, performing model training processing on the first set model based on a preset loss function and a preset training algorithm, and obtaining a first information recognition model; setting the label length parameter of the first information recognition model as a second value to obtain a second set model, inputting the training interaction type information data and the training interaction information corresponding to the training interaction type information data into the second set model, performing model training processing on the second set model based on the preset loss function and the preset training algorithm, and obtaining a trained information recognition model.

[0127] In the embodiment of the present disclosure, the training one-way type information data refers to a set of independent information samples that do not trigger any reply behavior, and each piece of data contains complete original information data content and a corresponding risk label. The training interaction type information data is a composite sample composed of initial information data and multiple rounds of dialogue triggered thereby, and the associated training interaction information contains complete dialogue thread data, i.e., full-chain interaction records from the initial information to the final reply.

[0128] In the data preparation phase, a dialogue tree pruning technique can be used to ensure data purity: for interactive data, only dialogue branches with direct causal relationship with the initial information are retained; for unidirectional data, potential long-delay reply samples are excluded through negative sample verification mechanism. Both types of data need to meet the size balance principle to prevent class weight deviation during model training.

[0129] The token length parameter is the core ability threshold of the information recognition model processing text sequence, and its value directly determines the upper limit of the text length that the model can analyze. During model training, a two-stage parameter setting strategy can be used. In the first setting model stage, the token length is set to a first value (e.g., 40-60 characters), which is optimized according to the length distribution characteristics of unidirectional information data to ensure that more than 95% of the single information content is covered. In the second setting model stage, it is adjusted to a second value (e.g., 150-250 characters), and the expansion capacity of the token length can be used to carry the multi-round dialogue content of interactive information. Parameter switching can be realized through model input layer reconstruction: when setting a new token length value, the sequence accommodation dimension of the word embedding matrix is dynamically adjusted, while the parameters of the trained feature extraction layer are retained.

[0130] The first stage training takes the first setting model as the carrier, and the input layer dimension is adapted to the token length of the first value. When the unidirectional information data is input into the model, the preset training algorithm (such as using the back propagation algorithm) is executed to make the model learn the semantic risk features of single information through iterative optimization. The preset loss function (such as cross-entropy loss) continuously monitors the deviation between the prediction results and the true labels to drive model parameter updates.

[0131] The second stage training takes the first information recognition model output by the previous stage as the basis, and expands its token length parameter to the second value to form the second setting model. At this time, the interactive information data and its interactive information are input, and the same preset loss function and training algorithm are used to learn the risk transmission pattern in the dialogue context. It should be noted that the two stages share the same underlying feature encoder, but generate output layer parameters adapted to different types of data, and finally fuse into a unified trained information recognition model.

[0132] Further, regarding the training method of the model, the following methods can also be used, but are not limited to:

[0133] Unidirectional abnormal short message recognition model training:

[0134] Use the pre-trained BERT model for training, and the training data includes all unidirectional short message content sent by the sender. Since the unidirectional short message only involves a single short message, the token length (tokenizer's max_length parameter) can be set smaller, such as 40.

[0135] During training, the content of the short message is input into the BERT model, and the model parameters are adjusted through the back propagation algorithm to minimize the loss function. The loss function uses the cross-entropy loss function.

[0136] Interaction type abnormal short message recognition model training:

[0137] The sender and each recipient are treated as a group, and the content of each pair of interaction short messages is trained. The training data includes all interaction short message content between the sender and each recipient. One implementation method is to concatenate all interaction short message content as a whole for training. Since the interaction type short message involves multiple short messages, the token length (tokenizer max_length parameter) can be set larger, such as: 200.

[0138] The pre-trained BERT model is used for training, and during the training process, the interaction short message content is input into the BERT model, and the model parameters are adjusted through the back propagation algorithm to minimize the loss function. The loss function uses the cross-entropy loss function.

[0139] Model evaluation:

[0140] The cross-validation method is used to evaluate the performance of the model, and the training data set is divided into training set and validation set. The accuracy, recall and F1 value of the model on the validation set are calculated, and the model parameters are adjusted according to the evaluation results to optimize the model performance.

[0141] Model optimization:

[0142] According to the evaluation results, the model parameters such as learning rate, batch size, etc. are adjusted to improve the recognition accuracy of the model. Regularization techniques (such as L1, L2 regularization) are used to prevent overfitting. Ensemble learning methods (such as: Bagging, Boosting) are used to improve the generalization ability of the model.

[0143] Further, it needs to be explained that when training the model, the model can be trained first through the interaction type information data, and then trained through the one-way type information data, or the model can be trained through the interaction type information data and the one-way type information data at the same time. The order of training the model through the interaction type information data and the one-way type information data is not limited by the embodiments of the present disclosure.

[0144] In an implementation manner of the embodiment of the present disclosure, after obtaining the information probability value, the following method can be used, but is not limited to: determining whether the information probability value is greater than a preset probability threshold; in a case where it is determined that the information probability value is not greater than the preset probability threshold, determining information data for which the information probability value is not greater than the preset probability threshold as normal information data; and in a case where it is determined that the information probability value is greater than the preset probability threshold, determining information data for which the information probability value is greater than the preset probability threshold as abnormal information data.

[0145] In the embodiment of the present disclosure, the preset probability threshold is a risk determination boundary value that is preset, and the value range of the preset probability threshold is set in the interval [0, 1] and is dynamically adjusted according to a security policy of an actual application scenario. When performing comparison, the information probability value output for each information data (that is, the risk confidence quantified by the model) is compared in real time.

[0146] When it is confirmed that the information probability value is not greater than the preset probability threshold, the corresponding information data is marked as normal information data. Further, when determining the information data, the information probability value needs to be strictly lower than the preset probability threshold, and no high-risk keywords (such as sensitive terms such as transfer and password) are detected. In the marking generation process, a determination basis chain is recorded synchronously, including but not limited to explainable data such as a probability value deviation degree and a key feature contribution degree. The normal information data will obtain a safe pass identification and will be excluded from a risk disposal queue in a subsequent processing process.

[0147] When it is confirmed that the information probability value is greater than the preset probability threshold, abnormal information data confirmation can be performed. First, a result review is performed, a feature contribution source of the high probability value is analyzed, then an association verification is performed, a risk consistency of other information in the same dialogue thread is scanned, and finally a disposal instruction set is generated, and a disposal level (such as early warning, interception, and traceability) is automatically matched according to a magnitude of the probability value exceeding the threshold. The finally generated abnormal mark can include three attributes: a risk level (divided according to a probability value interval), a risk type (mapped according to a feature of a model output layer), and a disposal suggestion (matched according to a predefined policy library).

[0148] Further, to cope with the characteristics of continuous evolution of fraud patterns, threshold dynamic optimization can be performed, verification feedback of determined normal information data and abnormal information data in an actual business scenario is continuously collected, when it is found that the misjudgment rate of a certain type rises, the threshold offset of the corresponding category is automatically adjusted. For example, when a new type of speech of a friend-making type abnormal message causes a false negative, the preset probability threshold of this category is reduced; and when a promotion type normal message is misjudged, the corresponding preset probability threshold is increased.

[0149] The dynamic comparison mechanism ensures that the risk judgment responds to business changes in real time, overcoming the lagging defects of static rule libraries; the binary processing of normal and abnormal information forms a clear disposal path, greatly improving the system execution efficiency; the threshold adaptive mechanism enables the system to have continuous evolution capability, effectively responding to the iterative updates of fraud means. Compared with the traditional binary output model, this confidence level-based judgment system significantly improves the fineness and interpretability of risk identification.

[0150] In an implementation manner of the embodiment of the present disclosure, before the data analysis and processing of the plurality of information data, the following methods can also be used, but are not limited to: performing value assignment processing on the plurality of signaling ends to obtain initial feature values corresponding to the plurality of signaling ends respectively; obtaining signaling end attribute data corresponding to the plurality of signaling ends respectively, and adjusting the initial feature values according to the signaling end attribute data and a preset risk condition to obtain target feature values corresponding to the plurality of signaling ends respectively, and determining a maximum feature value in the plurality of target feature values; performing data normalization processing on the plurality of target feature values and the maximum feature value respectively to obtain first risk scores corresponding to the plurality of signaling ends respectively; and determining the signaling end with the first risk score greater than a first preset risk threshold as a risk signaling end.

[0151] In the embodiment of the present disclosure, the initial feature value is a value set by the user, which is the starting value of risk assessment established for each signaling end. Regarding the setting of the initial feature value, a feature vector initialization strategy can be used to generate a feature vector containing N dimensions (N≥12) for each signaling end, and each dimension corresponds to a type of risk feature (such as age distribution, card opening behavior, etc. The initial feature value, i.e., the initial weight score of each dimension, is uniformly set to a reference value (such as 1) using the equal weight allocation principle.

[0152] The signaling end attribute data is a structured data set collected from multiple source systems such as communication metadata, user portrait library, and device fingerprint library, including but not limited to: static attributes (such as certificate type, registration channel) and dynamic behaviors (such as software usage record, social activity level).

[0153] The preset risk condition is a condition set by the user, including but not limited to: age, certificate ownership, whether to open cards in a concentrated manner, opening cards in 2 or more channels on the same day with the same certificate, repeatedly opening and closing cards with the same certificate in a short period of time, non-offline card opening, being stopped by other models, private chat software, remote control software, virtual transactions, etc.

[0154] Then the attribute data can be mapped to the preset risk dimension according to the preset risk condition, for example: the non-offline card opening attribute is mapped to the registration channel risk dimension, and the virtual transaction behavior is mapped to the fund movement dimension, etc.

[0155] The adjustment process is a dynamic correction process of the initial feature value based on preset risk conditions. When a preset risk condition is activated (e.g., detecting the same certificate on the same day in 2 or more channels), the initial feature value is multiplied by a preset coefficient (e.g., 2), and then the chain adjustment is started (e.g., further adjustment according to other preset risk conditions (repeated card spending)).

[0156] The data normalization process solves the horizontal comparability problem of different signaling terminal risk values. First, the maximum feature value in the total signaling terminal is determined as the reference denominator, and then the ratio of the target feature value of each signaling terminal to the maximum value is calculated to generate the first risk score in the interval [0, 1]. The calculation process can be optimized using a sliding window, such as only selecting the current active signaling terminal to participate in the maximum value calculation to avoid interference of historical frozen accounts with the reference value. The normalization result retains four decimal precision to ensure that the subtle differences of high-risk signaling terminals can still be identified.

[0157] When the first risk score of the signaling terminal exceeds the first preset risk threshold (e.g., 0.65-0.85), it is marked as a risk signaling terminal. The first preset risk threshold can be dynamically floating according to the real-time risk control strategy: automatically lowered in the high-fraud period to expand the monitoring range, and raised in the stable period to reduce false positives.

[0158] The construction of a dynamic evaluation system can at least achieve the following effects: initial equal weight assignment ensures the objectivity of the evaluation benchmark, laying the foundation for subsequent precise adjustment; intelligent mapping of attribute data to risk dimensions enables the system to adapt to new threat features; normalization based on the maximum feature value eliminates evaluation scale differences, making risk values of different magnitudes comparable. These features collectively improve the timeliness and accuracy of risk signaling terminal identification and reduce the probability of fraud information dissemination.

[0159] In an implementation manner of the embodiment of the present disclosure, after obtaining the information probability value, the following methods can be used, but are not limited to: obtaining a first information quantity corresponding to a plurality of information data, a second information quantity of cross-region information in the plurality of information data, and a third information quantity of reply information received by the risk signaling end; calculating and processing according to the first information quantity and the second information quantity to obtain a cross-region information proportion, and calculating and processing according to the first information quantity and the third information quantity to obtain an information transceiving ratio; performing data calculation and processing according to the first information quantity, the cross-region information proportion, and the information transceiving ratio to obtain a second risk score corresponding to the risk signaling end; and performing calculation and processing according to the first risk score, a preset weight, and the second risk score to obtain a comprehensive risk score corresponding to the risk signaling end, and performing shutdown processing on the risk signaling end in a case where the comprehensive risk score is greater than a second preset risk threshold.

[0160] In the embodiment of the present disclosure, the first information quantity refers to the total amount of information sent by the risk signaling end in a monitoring period, which is obtained by real-time statistics through a message flow log. The second information quantity specifically refers to the cumulative value of cross-region information in the information flow, and the determination standard of the cross-region information is that the receiving end belongs to a different provincial administrative division (which can be extended to a municipal level) from the signaling end registration place. The third information quantity refers to the total amount of valid reply information received by the risk signaling end, which excludes system automatic reply and port message.

[0161] The cross-region information proportion is the ratio of the second information quantity to the first information quantity, which reflects the geographical abnormality degree of information transmission. The information transceiving ratio is the ratio of the third information quantity to the first information quantity, which quantifies the information interaction activity.

[0162] The comprehensive risk score is jointly determined by the pre-portfolio risk (the first risk score) and the real-time behavior risk (the second risk score). The first risk score (pre-static evaluation result) and the second risk score (in-process dynamic behavior evaluation) are obtained, and linear fusion is performed through a preset weight:

[0163] Comprehensive risk score = α × first risk score + β × second risk score

[0164] The weight coefficients satisfy α + β = 1, and the weights are set according to the dynamic configuration of the risk control strategy (for example, β is increased to 0.7 in a high-fraud period, and α is increased to 0.6 in a daily period). The fusion result is normalized and mapped to a [0, 10] risk scale, which is convenient for cross-signaling end horizontal comparison.

[0165] When the comprehensive risk score exceeds the second preset risk threshold (a value set by the user, such as 7.5), a hierarchical shutdown protocol is triggered: primary shutdown: suspend the right to send new information, and the inventory information enters the audit queue; intermediate shutdown: block all information transmission channels and freeze the associated account; advanced shutdown: unregister the sender's identification, and report to the regulatory blacklist. Further, the decision-making process can set a dissent appeal channel, and the suspended sender can submit whitelist verification materials to start manual review.

[0166] By cross-regional information proportion accurate identification of geographical features; information transmission ratio effectively captures the behavior fingerprint of one-way induced abnormal short messages; the double-layer risk fusion model overcomes the dynamic lag defect of static portrait. It can identify new risk sources with normal pre-event characteristics but abnormal communication behavior, greatly improving the coverage and response speed of risk prevention and control.

[0167] In an implementable manner of an embodiment of the present disclosure, the calculation of the comprehensive risk score can use the following formula:

[0168]

[0169] Wherein, X is the comprehensive risk score, M i is the information probability value corresponding to the i-th information data in the plurality of information data, N is the first information quantity, W is the cross-regional information proportion, and F is the information transmission ratio.

[0170] In an implementable manner of an embodiment of the present disclosure, a system based on multi-dimensional information recognition is also provided, and the system architecture mainly consists of a data collection module, a data preprocessing module, a model training module, and a recognition module.

[0171] The data collection module is responsible for collecting all short messages sent by the sender and the corresponding reply short messages.

[0172] The data preprocessing module cleans and formats the collected data to ensure the consistency and accuracy of the data.

[0173] The model training module uses a pre-trained BERT model for training to generate an abnormal short message recognition model for different types. Further, other similar models can also be used, and the present disclosure is not limited thereto.

[0174] The recognition module inputs the short message to be detected into the trained model to output an abnormal information probability value and a type detection result.

[0175] Data preprocessing includes but is not limited to:

[0176] Data cleaning: remove irrelevant characters, punctuation marks and stop words in the short message, and keep the effective information in the short message. Formatting: convert the content of the short message into a unified format, which is convenient for subsequent model training and recognition.

[0177] In summary, the embodiments of the present disclosure can achieve the following technical effects:

[0178] 1. By analyzing all information or multiple information sent by the sender based on the behavior pattern of the sender, the category of the information is determined, that is, single-way information and interactive information. The single-way information is directly identified, and the interactive information is identified. By distinguishing single-way information and interactive information, and identifying the content of the information, and combining the information itself and the interactive information replied by the receiver, the accuracy of information identification is improved. The case that a single information is normal but multiple information can be determined as abnormal information can be effectively identified, thereby improving the accuracy and comprehensiveness of information identification.

[0179] 2. By comprehensively analyzing all short messages or multiple short messages sent by the sender, the behavior pattern of the sender can be more comprehensively understood, thereby improving the accuracy of abnormal short message identification.

[0180] 3. According to the division of abnormal short messages into single-way and interactive, the recognition model can be more accurately trained and recognized for different types of abnormal short messages.

[0181] 4. The pre-trained BERT model is used for training, combined with the cross-entropy loss function and the back propagation algorithm, which can effectively improve the recognition performance of the model.

[0182] 5. Not only based on the content of the short message, but also based on the sender's prior risk and the number of sent short messages to determine the comprehensive risk of the sender.

[0183] 6. It can be applied to telecommunication operators, security software providers, financial institutions and other scenarios to provide comprehensive short message security protection for users.

[0184] 7. Through the distributed architecture and containerization technology, the recognition model can be efficiently deployed and maintained, and the recognition speed and processing capacity can be improved.

[0185] Corresponding to the above-mentioned method for identifying information based on multiple dimensions, the present application also provides a device for identifying information based on multiple dimensions. Since the device embodiment of the present application corresponds to the method embodiment described above, for details not disclosed in the device embodiment, reference can be made to the method embodiment described above, which will not be described in detail in the present application.

[0186] Figure 2 The structure diagram of a device for identifying information based on multiple dimensions provided by the embodiments of the present disclosure is as follows: Figure 2As shown, comprising:

[0187] The analysis unit 21 is configured to perform data analysis processing on the identified multiple information data sent by the risk information sender, and determine the respective corresponding receiving terminals of the multiple information data.

[0188] The acquisition unit 22 is configured to acquire information interaction data between the risk information sender and the respective corresponding receiving terminals of the multiple information data.

[0189] The classification unit 23 is configured to perform classification processing on the multiple information data according to the information interaction data, to obtain unidirectional class information data and interactive class information data, wherein the unidirectional class information data is information data in the multiple information data without reply information, and the interactive class information data is information data in the multiple information data with reply information.

[0190] The recognition unit 24 is configured to input the unidirectional class information data and / or the interactive class information data and the corresponding interaction information of the interactive class information data into a trained information recognition model for information recognition processing, to obtain the respective corresponding information probability values of the multiple information data, wherein the trained information recognition model at least includes recognition parameters of the unidirectional class information data and recognition parameters of the interactive class information data.

[0191] The determination unit 25 is configured to determine information data with an information probability value greater than a preset probability threshold as abnormal information data.

[0192] Further, in a possible implementation manner of the embodiment of the present disclosure, the acquisition unit 22 is further configured to:

[0193] Determine a target receiving terminal that sends reply information to the risk information sender from the respective corresponding receiving terminals of the multiple information data;

[0194] Acquire information interaction data between the target receiving terminal and the risk information sender.

[0195] Further, in a possible implementation manner of the embodiment of the present disclosure, the classification unit 23 is further configured to:

[0196] Determine information data with reply information from the multiple information data according to the information interaction data;

[0197] Determine the information data with reply information as the interactive class information data, and acquire the corresponding interaction information of the interactive class information data from the information interaction data;

[0198] Determine information data without reply information as the unidirectional class information data.

[0199] Further, in a possible implementation of the embodiment of the present disclosure, the obtaining unit 22 is further configured to obtain one-way class information data for training, interaction class information data for training, and interaction information for training corresponding to the interaction class information data for training;

[0200] As shown in Figure 3 the device for identifying information based on multiple dimensions further comprises:

[0201] The training unit 26 is configured to set a mark length parameter of the to-be-trained model as a first value to obtain a first set model, input the one-way class information data for training into the first set model, and perform model training processing on the first set model based on a preset loss function and a preset training algorithm to obtain a first information identification model.

[0202] The training unit 26 is further configured to set the mark length parameter of the first information identification model as a second value to obtain a second set model, input the interaction class information data for training and the interaction information for training corresponding to the interaction class information data for training into the second set model, and perform model training processing on the second set model based on the preset loss function and the preset training algorithm to obtain a trained information identification model.

[0203] Further, in a possible implementation of the embodiment of the present disclosure, the determining unit 25 is further configured to:

[0204] determine whether the information probability value is greater than a preset probability threshold;

[0205] in a case where it is determined that the information probability value is not greater than the preset probability threshold, determine information data for which the information probability value is not greater than the preset probability threshold as normal information data;

[0206] in a case where it is determined that the information probability value is greater than the preset probability threshold, determine information data for which the information probability value is greater than the preset probability threshold as abnormal information data.

[0207] Further, in a possible implementation of the embodiment of the present disclosure, as shown in Figure 3 the device for identifying information based on multiple dimensions further comprises:

[0208] The processing unit 27 is configured to perform value assignment processing on the multiple signaling ends to obtain initial feature values corresponding to the multiple signaling ends respectively;

[0209] The processing unit 27 is further configured to obtain signaling end attribute data corresponding to the multiple signaling ends respectively, and perform adjustment processing on the initial feature values according to the signaling end attribute data and a preset risk condition to obtain target feature values corresponding to the multiple signaling ends respectively, and determine a maximum feature value in the multiple target feature values;

[0210] The processing unit 27 is further configured to perform data normalization on the plurality of target feature values and the maximum feature value respectively to obtain a first risk score corresponding to each of the plurality of information sending terminals respectively.

[0211] The processing unit 27 is further configured to determine the information sending terminal with the first risk score greater than a first preset risk threshold as a risk information sending terminal.

[0212] Further, in a possible implementation of the embodiment of the present disclosure, the obtaining unit 22 is further configured to obtain a first information quantity corresponding to the plurality of information data, a second information quantity of cross-region information in the plurality of information data, and a third information quantity of reply information received by the risk information sending terminal.

[0213] The processing unit 27 is further configured to perform calculation processing on the first information quantity and the second information quantity to obtain a cross-region information proportion, and perform calculation processing on the first information quantity and the third information quantity to obtain an information transceiving ratio.

[0214] The processing unit 27 is further configured to perform data calculation processing on the first information quantity, the cross-region information proportion, and the information transceiving ratio to obtain a second risk score corresponding to the risk information sending terminal.

[0215] The processing unit 27 is further configured to perform calculation processing on the first risk score, a preset weight, and the second risk score to obtain a comprehensive risk score corresponding to the risk information sending terminal, and perform shutdown processing on the risk information sending terminal in a case where the comprehensive risk score is greater than a second preset risk threshold.

[0216] Further, in a possible implementation of the embodiment of the present disclosure, the processing unit is further configured to:

[0217] The comprehensive risk score is calculated by using the following formula:

[0218]

[0219] wherein, X is the comprehensive risk score, M i is an information probability value corresponding to the i-th information data in the plurality of information data, N is the first information quantity, W is the cross-region information proportion, and F is the information transceiving ratio.

[0220] It should be noted that the above explanation and description of the method embodiment are also applicable to the device of the embodiment of the present disclosure, and the principle is the same, which is not limited in the embodiment of the present disclosure.

[0221] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0222] Figure 4A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0223] like Figure 4 As shown, device 400 includes a computing unit 401, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 402 or a computer program loaded from storage unit 408 into RAM (Random Access Memory) 403. RAM 403 may also store various programs and data required for the operation of device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via bus 404. I / O (Input / Output) interface 405 is also connected to bus 404.

[0224] Multiple components in device 400 are connected to I / O interface 405, including: input unit 406, such as keyboard, mouse, etc.; output unit 407, such as various types of monitors, speakers, etc.; storage unit 408, such as disk, optical disk, etc.; and communication unit 409, such as network card, modem, wireless transceiver, etc. Communication unit 409 allows device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0225] The computing unit 401 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 performs various methods and processes described above, such as the method of multi-dimension based information recognition. For example, in some embodiments, the method of multi-dimension based information recognition can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded onto the RAM 403 and executed by the computing unit 401, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to perform the aforementioned method of multi-dimension based information recognition by any other appropriate means, such as by means of firmware.

[0226] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a Field Programmable Gate Array (FPGA), an Application-Specific Integrated Circuit (ASIC), an Application Specific Standard Product (ASSP), a System on Chip (SOC), a Complex Programmable Logic Device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0227] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0228] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable storage medium can include but are not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include one or more lines of electrical wire, portable computer diskette, hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory), or flash memory, fiber optics, CD-ROM (Compact Disc Read-Only Memory), optical storage device, magnetic storage device, or any suitable combination of the foregoing.

[0229] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (Cathode Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0230] The systems and techniques described here can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.

[0231] The computer system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server is one of communication and distribution, with the server receiving requests from the client and transmitting responses via the communication network. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system. The server can also be a server of a distributed system, or a server combined with a blockchain.

[0232] It should be noted that artificial intelligence is a discipline that studies enabling computers to simulate some thinking processes and intelligent behaviors of people (such as learning, reasoning, thinking, planning, etc.), which has both hardware and software technologies. Artificial intelligence hardware technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing, etc.; artificial intelligence software technology mainly includes computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, knowledge graph technology, etc. several major directions.

[0233] It should be understood that various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which is not limited herein.

[0234] The above detailed description does not limit the scope of the disclosure. Various modifications, combinations, sub-combinations and alternatives can be made to the detailed description. Any modification, equivalent replacement and improvement etc. made within the spirit and principle of the disclosure shall be included in the scope of the disclosure.

Claims

1. A method for multi-dimensional information recognition, characterized by, The method comprises the following steps: performing data analysis processing on the multiple information data sent by the identified risk sender, determining the respective corresponding receivers of the multiple information data, and obtaining information interaction data between the risk sender and the respective corresponding receivers of the multiple information data; performing classification processing on the multiple information data according to the information interaction data, obtaining unidirectional class information data and interactive class information data, wherein the unidirectional class information data is information data without reply information in the multiple information data, and the interactive class information data is information data with reply information in the multiple information data; inputting the unidirectional class information data and / or the interactive class information data and the corresponding interactive information of the interactive class information data into a trained information recognition model to perform information recognition processing, and obtaining information probability values corresponding to the multiple information data, wherein the trained information recognition model at least includes recognition parameters of the unidirectional class information data and recognition parameters of the interactive class information data; information data with an information probability value greater than a preset probability threshold is determined as abnormal information data.

2. The method of claim 1, wherein, The obtaining of the information interaction data between the risk sender and the respective corresponding receivers of the multiple information data comprises: determining a target receiver that sends the reply information to the risk sender among the respective corresponding receivers of the multiple information data; obtaining the information interaction data between the target receiver and the risk sender.

3. The method of claim 2, wherein, The classification processing of the multiple information data according to the information interaction data to obtain unidirectional class information data and interactive class information data comprises: determining information data with the reply information in the multiple information data according to the information interaction data; determining information data with the reply information as the interactive class information data, and obtaining the corresponding interactive information of the interactive class information data from the information interaction data; determining information data without the reply information as the unidirectional class information data.

4. The method of claim 1, wherein, Before the data analysis processing on the multiple information data sent by the identified risk sender to determine the respective corresponding receivers of the multiple information data, the method further comprises: obtaining training unidirectional class information data, training interactive class information data, and corresponding training interactive information of the training interactive class information data; setting a mark length parameter of a to-be-trained model as a first value to obtain a first set model, inputting the training unidirectional class information data into the first set model, performing model training processing on the first set model based on a preset loss function and a preset training algorithm, and obtaining a first information recognition model; setting a mark length parameter of the first information recognition model as a second value to obtain a second set model, inputting the training interactive class information data and the corresponding training interactive information of the training interactive class information data into the second set model, performing model training processing on the second set model based on the preset loss function and the preset training algorithm, and obtaining the trained information recognition model.

5. The method for multi-dimension based information recognition according to claim 1, wherein, After the one-way information data, and / or the interactive information data and the corresponding interactive information are input into the trained information recognition model for information recognition processing, the method further comprises: determining whether the information probability value is greater than the preset probability threshold value; in the case where it is determined that the information probability value is not greater than the preset probability threshold value, determining the information data whose information probability value is not greater than the preset probability threshold value as normal information data; in the case where it is determined that the information probability value is greater than the preset probability threshold value, determining the information data whose information probability value is greater than the preset probability threshold value as the abnormal information data.

6. The method for multi-dimension based information recognition according to claim 1, wherein, Before the data analysis processing of the identified multiple information data sent by the risk sender and the determination of the corresponding receiving end of the multiple information data, the method further comprises: performing value assignment processing on multiple sending ends to obtain the initial feature value corresponding to each of the multiple sending ends; obtaining the sending end attribute data corresponding to each of the multiple sending ends, and adjusting the initial feature value according to the sending end attribute data and the preset risk condition to obtain the target feature value corresponding to each of the multiple sending ends, and determining the maximum feature value in the multiple target feature values; performing data normalization processing on the multiple target feature values and the maximum feature value respectively to obtain the first risk score corresponding to each of the multiple sending ends; determining the sending end whose first risk score is greater than a first preset risk threshold as the risk sender.

7. The method of claim 6, wherein, After the one-way information data, and / or the interactive information data and the corresponding interactive information are input into the trained information recognition model for information recognition processing, the method further comprises: obtaining the first information quantity corresponding to the multiple information data, the second information quantity of cross-regional information in the multiple information data, and the third information quantity of reply information received by the risk sender; performing calculation processing according to the first information quantity and the second information quantity to obtain a cross-regional information proportion, and performing calculation processing according to the first information quantity and the third information quantity to obtain an information transceiving ratio; performing data calculation processing according to the first information quantity, the cross-regional information proportion, and the information transceiving ratio to obtain a second risk score corresponding to the risk sender; performing calculation processing according to the first risk score, a preset weight, and the second risk score to obtain a comprehensive risk score corresponding to the risk sender, and performing shutdown processing on the risk sender in the case where the comprehensive risk score is greater than a second preset risk threshold.

8. The method of claim 7, wherein, The calculation processing according to the first risk score, a preset weight, and the second risk score to obtain the comprehensive risk score corresponding to the risk sender comprises: the calculation of the comprehensive risk score adopts the following formula: wherein X is the comprehensive risk score, M i is the information probability value corresponding to the i-th information data in the plurality of information data, N is the first information quantity, W is the cross-region information proportion, and F is the information sending-receiving ratio.

9. A device for multi-dimensional information recognition, characterized in that, comprises: An analysis unit is configured to perform data analysis on the identified information data from the risk sender to determine a corresponding receiver of each of the information data; An acquisition unit is configured to acquire information interaction data between the risk sender and the corresponding receiver of each of the information data; A classification unit is configured to classify the information data according to the information interaction data to obtain unidirectional class information data and interactive class information data, wherein the unidirectional class information data is information data without reply information in the information data, and the interactive class information data is information data with reply information in the information data; An identification unit is configured to input the unidirectional class information data and / or the interactive class information data and corresponding interactive information of the interactive class information data into a trained information identification model to perform information identification processing to obtain an information probability value corresponding to each of the information data, wherein the trained information identification model at least includes identification parameters of the unidirectional class information data and identification parameters of the interactive class information data; A determination unit is configured to determine information data with an information probability value greater than a preset probability threshold as abnormal information data.

10. An electronic device, comprising: comprise: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.

Citation Information

Cited By

  • Natural scene text detection method based on semantic feature step-by-step recombination

    CN121789197A

  • A natural scene text detection method based on semantic feature step-by-step reorganization

    CN121789197B