Method and apparatus for identifying network data risk, and program product
Through multiple optimizations of Netxun data and the application of the bidirectional encoder representation model, the problem that the model cannot self-optimize in Netxun data risk identification is solved, and refined risk identification and intelligent audit are achieved, which improves the accuracy and efficiency of risk identification.
Patent Information
- Application Number
- CN202510628569.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, traditional text analysis and rule-based methods lack online learning and instant feedback mechanisms when identifying the risks of online data, resulting in the inability to self-optimize the model and inaccurate risk identification results.
The historical network data is calculated through the Nth optimization target evaluation model, combined with the gradient descent algorithm to optimize the model, and a bidirectional encoder representation model is used to build a target knowledge graph, and input the N+1 optimization model for risk assessment.
It realizes refined risk identification, enhances the adaptability and predictive stability of the model, improves audit efficiency and information dissemination security, and supports intelligent auditing and immediate response.
Smart Images

Figure CN120338511A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular, to a method, apparatus, and program product for identifying risks in network communication data. Background Art
[0002] In the current field of bank information management, especially in the review process of network communication, the challenges are becoming increasingly prominent. The traditional method relying on manual review has many deficiencies, such as low review efficiency, difficulty in ensuring the consistency of review results, and limited ability to identify complex associations and implicit risks between information. In addition, existing technologies generally lack the ability to dynamically adjust review strategies and the rapid adaptability to new threats and compliance requirements.
[0003] At the same time, although there have been various attempts to improve this situation, including using traditional text analysis and rule-based methods, these solutions have obvious limitations in dealing with complex network structure data and capturing deep-level associations between information. Although the early graph neural network technology introduced a network perspective, it ignored the two-way nature of information propagation, reducing the comprehensiveness and accuracy of the model. More critically, these methods generally lack an online learning and instant feedback mechanism, making the model unable to self-optimize with changes in the environment, and greatly reducing the long-term effectiveness.
[0004] Regarding the problem that when using traditional text analysis and rule-based methods to identify risks in network communication data in related technologies, the lack of an online learning and instant feedback mechanism makes the model unable to self-optimize, resulting in inaccurate risk identification results, no effective solution has been proposed yet. Summary of the Invention
[0005] The main objective of this application is to provide a method, apparatus, and program product for identifying risks in network communication data, so as to solve the problem that when using traditional text analysis and rule-based methods to identify risks in network communication data in related technologies, the lack of an online learning and instant feedback mechanism makes the model unable to self-optimize, resulting in inaccurate risk identification results.
[0006] To achieve the above object, according to one aspect of the present application, a method for identifying risks of network communication data is provided. The method includes: calculating a first risk score for a historical knowledge graph corresponding to historical network communication data through an Nth optimized target evaluation model, and determining historical risk data based on the first risk score, where N is a non-negative integer; optimizing the Nth optimized target evaluation model based on the audit result of the historical risk data according to the gradient descent algorithm to obtain an (N + 1)th optimized target evaluation model; processing the currently to-be-processed network communication data using a bidirectional encoder representation model to construct a target knowledge graph; inputting the target knowledge graph into the (N + 1)th optimized target evaluation model, calculating the risk score of the to-be-processed network communication data, and generating a risk assessment result.
[0007] Further, processing the currently to-be-processed network communication data using a bidirectional encoder representation model to construct a target knowledge graph includes: performing data preprocessing on the to-be-processed network communication data to obtain preprocessed data; using the bidirectional encoder representation model to extract features from the preprocessed data to obtain a high-dimensional feature representation; performing named entity recognition and relationship extraction on the high-dimensional feature representation based on a transformer model to obtain multiple entities and the relationships between the multiple entities; using a graph structure to represent the multiple entities and the relationships between the multiple entities to obtain the target knowledge graph.
[0008] Further, performing data preprocessing on the to-be-processed network communication data to obtain preprocessed data includes: using a regular expression to clean the to-be-processed network communication data, performing word segmentation on the cleaned network communication data, and converting the segmented data into input sequence data, where the input sequence data is in the format of the input data of the bidirectional encoder representation model; adding a first marker at the start position of the input sequence data and a second marker at the end position of the input sequence data to obtain marked input sequence data; performing a data padding operation or a data truncation operation on the marked input sequence data according to a preset data length to obtain the preprocessed data.
[0009] Further, using the bidirectional encoder representation model to extract features from the preprocessed data to obtain a high-dimensional feature representation includes: converting the preprocessed data into an embedding representation; using the self-attention mechanism in the bidirectional encoder representation model to calculate the similarity between each marker in the input sequence corresponding to the embedding representation, and performing a weighted sum on the similarity to obtain the attention information of the embedding representation; using multiple encoder layers in the bidirectional encoder representation model to determine the complex context relationship of the embedding representation based on the attention information data of the embedding representation to generate the high-dimensional feature representation.
[0010] Further, converting the preprocessed data into an embedding representation includes: converting each token in the preprocessed data into a vector representation of a fixed size to obtain token embedding vectors; converting each segment in the preprocessed data into a vector representation to obtain segment embedding vectors; determining position embedding vectors according to the positions of each token in the input sequence in the preprocessed data; and determining the embedding representation according to the token embedding vectors, the segment embedding vectors, and the position embedding vectors.
[0011] Further, performing named entity recognition and relation extraction on the high-dimensional feature representation based on a Transformer model to obtain multiple entities and the relationships between the multiple entities, including: sequentially inputting the high-dimensional feature representation into the encoder and decoder of the Transformer model to output the entity labels of each token, where the entity labels are used to identify different types of entities; labeling the entity corresponding to each token by using a token algorithm based on the entity type of each token to obtain the entity annotation information of each token, and determining the multiple entities according to the entity annotation information of each token; identifying entity pairs with an association relationship among the multiple entities according to the entity annotation information of each token, and extracting the context information of each entity by using the bidirectional encoder representation model; and inputting each entity and the context information of each entity into the Transformer model to classify the relationships of the entity pairs to obtain the relationships between the multiple entities, where each relationship among the relationships between the multiple entities includes at least one of the following: cooperation relationship, subordination relationship, and time relationship.
[0012] Further, representing the multiple entities and the relationships between the multiple entities by using a graph structure to obtain the target knowledge graph, including: creating corresponding nodes in the graph according to each entity among the multiple entities; creating corresponding edges in the graph according to the entity pairs with relationships in the entity pairs; creating an initial knowledge graph according to the nodes and the edges, and constructing an adjacency matrix according to the initial knowledge graph; optimizing the nodes and edges included in the initial knowledge graph, updating the adjacency matrix, determining the target knowledge graph according to the updated adjacency matrix, and storing the target knowledge graph in a preset storage space.
[0013] Further, based on the gradient descent algorithm, the target evaluation model after the Nth optimization is optimized according to the review result of the historical risk data to obtain the target evaluation model after the (N + 1)th optimization, including: receiving the review result of the target object for the historical risk data, and annotating the review result to obtain the annotation information of the review result; performing data preprocessing on the review result and the annotation information of the review result to obtain the preprocessed review information; based on the gradient descent algorithm, using the preprocessed review information to optimize the model parameters of the target evaluation model after the Nth optimization to obtain the target evaluation model after the (N + 1)th optimization.
[0014] To achieve the above object, according to another aspect of the present application, there is provided a device for identifying the risk of network communication data, the device includes: a calculation unit, configured to calculate a first risk score for a historical knowledge graph corresponding to historical network communication data through a target evaluation model after the Nth optimization, and determine historical risk data according to the first risk score, where N is a non-negative integer; an optimization unit, configured to optimize the target evaluation model after the Nth optimization based on the gradient descent algorithm according to the review result of the historical risk data to obtain the target evaluation model after the (N + 1)th optimization; a construction unit, configured to process the currently to-be-processed network communication data by using a bidirectional encoder representation model to construct a target knowledge graph; a generation unit, configured to input the target knowledge graph into the target evaluation model after the (N + 1)th optimization, calculate the risk score of the to-be-processed network communication data, and generate a risk assessment result.
[0015] Further, the construction unit includes: a first processing subunit, configured to perform data preprocessing on the to-be-processed network communication data to obtain preprocessed data; an extraction subunit, configured to extract features from the preprocessed data by using the bidirectional encoder representation model to obtain a high-dimensional feature representation; an identification subunit, configured to perform named entity recognition and relationship extraction on the high-dimensional feature representation based on a transformer model to obtain a plurality of entities and the relationships between the plurality of entities; a characterization subunit, configured to characterize the plurality of entities and the relationships between the plurality of entities by using a graph structure to obtain the target knowledge graph.
[0016] Further, the processing subunit includes: a cleaning module, configured to perform data cleaning on the to-be-processed network communication data by using a regular expression, perform word segmentation on the cleaned network communication data, and convert the segmented data into input sequence data, where the input sequence data is in the format of the input data of the bidirectional encoder representation model; an adding module, configured to add a first marker at the start position of the input sequence data and add a second marker at the end position of the input sequence data to obtain the marked input sequence data; a padding module, configured to perform a data padding operation or a data truncation operation on the marked input sequence data according to a preset data length to obtain the preprocessed data.
[0017] Further, the extraction subunit includes: a conversion module, configured to convert the preprocessed data into an embedded representation; a calculation module, configured to calculate the similarity between each marker in the input sequence corresponding to the embedded representation by using the self-attention mechanism in the bidirectional encoder representation model, and perform a weighted sum on the similarity to obtain the attention information of the embedded representation; a generation module, configured to use multiple encoder layers in the bidirectional encoder representation model to determine the complex context relationship of the embedded representation according to the attention information data of the embedded representation and generate the high-dimensional feature representation.
[0018] Further, the conversion module includes: a first conversion sub-module, configured to convert each marker in the preprocessed data into a vector representation of a fixed size to obtain a marker embedding vector; a second conversion sub-module, configured to convert each segment in the preprocessed data into a vector representation to obtain a segment embedding vector; a first determination sub-module, configured to determine a position embedding vector according to the position of each marker in the input sequence in the preprocessed data; a second determination sub-module, configured to determine the embedded representation according to the marker embedding vector, the segment embedding vector, and the position embedding vector.
[0019] Further, the recognition subunit includes: an output module, configured to sequentially input the high-dimensional feature representation into the encoder and decoder of the transformer model, and output an entity label for each token, where the entity label is used to identify different types of entities; a labeling module, configured to label the entity corresponding to each token using a labeling algorithm based on the entity type of each token, obtain entity annotation information for each token, and determine the multiple entities according to the entity annotation information for each token; a recognition module, configured to recognize entity pairs having an association relationship among the multiple entities according to the entity annotation information for each token, and extract context information of each entity using the bidirectional encoder representation model; a classification module, configured to input each entity and the context information of each entity into the transformer model, classify the relationship of the entity pair, and obtain the relationship among the multiple entities, where each relationship among the multiple entities includes at least one of the following: a cooperation relationship, a subordination relationship, and a temporal relationship.
[0020] Further, the characterization subunit includes: a first creation module, configured to create a corresponding node in the graph for each entity among the multiple entities; a second creation module, configured to create a corresponding edge in the graph for the entity pairs having a relationship among the entity pairs; a third creation module, configured to create an initial knowledge graph according to the nodes and the edges, and construct an adjacency matrix according to the initial knowledge graph; a storage module, configured to optimize the nodes and edges included in the initial knowledge graph, update the adjacency matrix, determine the target knowledge graph according to the updated adjacency matrix, and store the target knowledge graph in a preset storage space.
[0021] Further, the optimization unit includes: a labeling subunit, configured to receive an audit result of the target object for the historical risk data, and label the audit result to obtain labeled information of the audit result; a second processing subunit, configured to perform data preprocessing on the audit result and the labeled information of the audit result to obtain preprocessed audit information; an optimization subunit, configured to optimize model parameters of the Nth optimized target evaluation model using the preprocessed audit information based on a gradient descent algorithm to obtain the (N + 1)th optimized target evaluation model.
[0022] To achieve the above object, according to one aspect of the present application, there is provided a computer program product, including a computer program, where when the computer program is executed by a processor, it implements the method for identifying network communication data risks described in any one of the above, and when the computer program is executed by a processor, it implements the steps of the method for identifying network communication data risks in each embodiment of the present application.
[0023] To achieve the above object, according to one aspect of the present application, there is provided a computer-readable storage medium, which includes stored computer instructions, wherein when the computer instructions are executed by a processor, the method for identifying the risk of network communication data described in any one of the above is implemented.
[0024] To achieve the above object, according to one aspect of the present application, there is provided an electronic device, including one or more processors and a memory, the memory is used to store one or more programs, wherein when one or more programs are executed by one or more processors, one or more processors are caused to implement the method for identifying the risk of network communication data described in any one of the above.
[0025] In the embodiment of the present application, the target evaluation model optimized for the Nth time is used to calculate the historical knowledge graph corresponding to the historical network communication data, and a first risk score is obtained, and the historical risk data is determined based on the first risk score, where N is a non-negative integer; based on the gradient descent algorithm, the target evaluation model optimized for the Nth time is optimized according to the review result of the historical risk data to obtain the target evaluation model optimized for the (N + 1)th time; the bidirectional encoder representation model is used to process the currently to-be-processed network communication data to construct a target knowledge graph; the target knowledge graph is input into the target evaluation model optimized for the (N + 1)th time, the risk score of the to-be-processed network communication data is calculated, and a risk assessment result is generated, thereby solving the technical problem that when the risk of network communication data is identified by using traditional text analysis and rule-based methods, the lack of an online learning and instant feedback mechanism makes the model unable to self-optimize, resulting in inaccurate risk identification results.
[0026] By using the target evaluation model optimized for the Nth time to calculate the historical knowledge graph corresponding to the historical network communication data, the risk level of the historical content can be accurately quantified, achieving the technical effect of refined risk identification. At the same time, through the (N + 1)th iterative optimization based on the gradient descent algorithm according to the review result of the historical risk data, the model parameters can be finely tuned, realizing the self-correction and efficiency improvement of the model, and further achieving the technical effect of enhancing the adaptability and prediction stability of the model; by using the bidirectional encoder representation model to process the currently to-be-processed network communication data to construct a target knowledge graph, the entities and relationships in the new information can be deeply mined and structured, achieving the technical effect of intelligent parsing and efficient information integration. Inputting this knowledge graph into the target evaluation model optimized for the (N + 1)th time, calculating the risk score of the to-be-processed network communication data, and generating a risk assessment result can real-time evaluate the content risk, realizing intelligent review and instant response, and further achieving the technical effect of improving the review efficiency and ensuring the security of information dissemination. Description of the Drawings
[0027] The accompanying drawings, which form a part of this application, are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0028] Figure 1 is a hardware structural block diagram of a computer terminal (or mobile device) for implementing a method for identifying risks of online communication data according to Embodiment 1 of this application;
[0029] Figure 2 is a flowchart of an optional method for identifying risks of online communication data according to Embodiment 1 of this application;
[0030] Figure 3 is a schematic diagram of an optional method for identifying risks of online communication data according to Embodiment 1 of this application;
[0031] Figure 4 is a schematic diagram of a device for identifying risks of online communication data according to Embodiment 2 of this application;
[0032] Figure 5 is a schematic diagram of an electronic device for identifying risks of online communication data according to Embodiment 3 of this application. Detailed implementation manners
[0033] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The following will refer to the accompanying drawings and combine with the embodiments to detail this application.
[0034] It should be noted that the processing methods, devices, storage media, and methods and devices for determining electronic devices in this application document can be used in the field of fintech to improve the accuracy of risk identification results during the process of identifying risks of online communication data, and can also be used in any field other than the field of fintech. The application fields of the methods and devices of the processing methods, devices, storage media, and electronic devices in this application document are not limited.
[0035] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, collected data, used data, generated data, processed data, etc.) and data (including but not limited to data for analysis, stored data, displayed data, collected information, used information, generated information, processed information, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, complies with the relevant laws, regulations, and standards of relevant countries and regions, adopts necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or reject. For example, there are interfaces set between this system and relevant users or institutions to provide corresponding operation entrances for users to choose to agree or reject the automated decision-making results; if the user chooses to reject, it will enter the expert decision-making process.
[0036] Embodiment 1
[0037] According to an embodiment of the present application, there is also provided a method embodiment for identifying the risk of network communication data. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from that here.
[0038] The method embodiment provided by Embodiment 1 of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or mobile device) for implementing the method for identifying the risk of network communication data is shown. As Figure 1 shown, the computer terminal 10 (or mobile device) may include one or more (shown as 102a, 102 for identifying the risk of network communication data,..., 102n in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the USB bus for identifying the risk of network communication data), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0039] It should be noted that one or more of the above-mentioned processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0040] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the method for identifying the risk of network communication data in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned method for identifying the risk of network communication data. The memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0041] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0042] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the computer terminal 10 (or mobile device).
[0043] Under the above operating environment, the present application provides a method for identifying the risk of network communication data as shown in Figure 2 the following. Figure 2 It is a flowchart of an optional method for identifying the risk of network communication data provided according to Embodiment 1 of the present application.
[0044] Step S201, calculate the historical knowledge graph corresponding to the historical network information data through the target evaluation model optimized for the Nth time to obtain a first risk score, and determine the historical risk data based on the first risk score, where N is a non-negative integer.
[0045] In this embodiment 1, the historical knowledge graph constructed by the historical web news data is deeply analyzed through the target evaluation model after multiple iterations of optimization, and the first risk score of each historical web news data is calculated. The first risk score quantifies the potential violation or security risk level of the historical web news data. Based on the first risk score, part of the content in the historical web news data is classified as historical risk data.
[0046] The target assessment model is a model obtained after training a bidirectional graph neural network (Bi-GNN for short). The model conducts in-depth analysis on each piece of online news data, taking into account multiple features in the content and their interrelationships. After Bi-GNN processing, a risk score p is output for each piece of content. {i} (For example, the first risk score mentioned above), which reflects the potential risk level of the content. This value ranges from 0 to 1, representing the probability of risk from low to high. The Sigmoid activation function (σ) is used to map the node feature h in the knowledge graph. {i} With the preset weight W {r} and bias b {r} The linear combination of can be expressed as: {i} =σ(W {r} ·h {i} +b {r} ), where h {i} represents the feature vector of node i, W {r} and b {r} They are the model parameters learned by Bi-GNN.
[0047] Furthermore, based on the risk score output by the target assessment model, a decision threshold (θ) (e.g., 0.5) is set to distinguish content of different risk levels. i >θ), it is considered high risk and requires a more stringent manual review process to ensure the safety and compliance of the content. i ≤θ), indicating that the content risk is low, and an automated release process is implemented to quickly approve the release and improve processing efficiency. This strategy achieves accurate classification and corresponding processing of different network information data through flexible adjustment of thresholds and automated risk assessment, which not only ensures content security but also optimizes the allocation of audit resources.
[0048] Furthermore, when a piece of online news data is initially determined to be of high risk, in order to ensure the accuracy and security of the information, a manual review process will be involved. Professionals will conduct a comprehensive review and analysis of the content of the online news. The task is to verify the information in the online news data, evaluate the authenticity of its content and the possible risks it may bring, and ensure that all the information published in the online news data is accurate and complies with relevant regulations.
[0049] Meanwhile, if the content in the online news data involves specific sensitive information, corresponding security protocols will be initiated according to its sensitivity level. The purpose of these security protocols is to further protect the security of the information and prevent the improper dissemination or abuse of sensitive information. This may include encrypting the information, restricting the access and dissemination scope of the information, or even modifying or deleting the information to ensure the legal and compliant use of the information. Meanwhile, if the content of the online news involves specific sensitive information, corresponding security protocols will be initiated according to its sensitivity level. The purpose of these security protocols is to further protect the security of the information and prevent the improper dissemination or abuse of sensitive information. This may include encrypting the information, restricting the access and dissemination scope of the information, or even modifying or deleting the information to ensure the legal and compliant use of the information.
[0050] Step S202: Optimize the target evaluation model optimized for the Nth time based on the historical risk data according to the gradient descent algorithm to obtain the target evaluation model optimized for the (N + 1)th time.
[0051] In this Embodiment 1, after the Nth optimization, the target evaluation model has been trained multiple times based on historical data, forming a preliminary understanding of the risky online news data. Next, by analyzing the review results of the previous historical risk data, the target evaluation model optimized for the Nth time can be iteratively optimized according to the online news data marked as high risk or low risk. In each iteration (i.e., from the Nth to the (N + 1)th optimization), the model adjusts the parameters according to the gradient and gradually approaches a more optimized state. After continuous training and optimization, the finally obtained model optimized for the (N + 1)th time will be more accurate, able to more reliably evaluate the risks of new online news content, and improve the efficiency and security of the review process. This process reflects the self-improving characteristic of the machine learning model through the feedback mechanism.
[0052] Step S203: Process the currently to-be-processed online news data using a bidirectional encoder representation model to construct a target knowledge graph.
[0053] In this Embodiment 1, the Bidirectional Encoder Representations from Transformers (which can be abbreviated as Bi-Encoder, or simply referred to as the BERT model) processes the currently received network news data to construct a target knowledge graph. Specifically, the Bi-Encoder model encodes each independent part of the network news (such as the title, text, keywords, etc.) separately and converts them into high-dimensional vector representations. These vectors can capture the semantic information of the text. At the same time, through bidirectional propagation, the model processes the text not only from front to back but also from back to front to ensure a complete understanding of the context.
[0054] Step S204: Input the target knowledge graph into the target evaluation model optimized for the (N + 1)-th time, calculate the risk score of the network news data to be processed, and generate a risk assessment result.
[0055] In this Embodiment 1, the constructed target knowledge graph is used as input and passed to the evaluation model optimized for the (N + 1)-th time to quantitatively analyze the risk level contained in the current network news data. The model calculates a risk score by analyzing the complex relationships and attributes between entities in the knowledge graph and combining the risk identification patterns learned during the previous training. This score intuitively reflects the degree of potential violations or security threats in the currently received network news data and provides a key basis for generating the subsequent risk assessment result. Network news with a high score will be regarded as high-risk and may trigger further manual review or other security measures; while network news with a low score can be regarded as relatively safe and may not require any processing.
[0056] Optionally, in the method for identifying the risk of network news data provided in Embodiment 1 of this application, a Bidirectional Encoder Representations model is used to process the currently to-be-processed network news data to construct a target knowledge graph, including: performing data preprocessing on the to-be-processed network news data to obtain preprocessed data; using the Bidirectional Encoder Representations model to extract features from the preprocessed data to obtain high-dimensional feature representations; performing named entity recognition and relationship extraction on the high-dimensional feature representations based on a Transformer model to obtain multiple entities and the relationships between multiple entities; using a graph structure to represent the multiple entities and the relationships between multiple entities to obtain a target knowledge graph.
[0057] In this Embodiment 1, the preprocessing operation includes cleaning and formatting the original network news data to ensure that the text is clean and standardized. For example, removing HTML tags, special characters, unifying the text format, and performing preprocessing operations such as word segmentation and tokenization, making the data suitable for the input requirements of the deep learning model, and at the same time improving the efficiency and accuracy of model training.
[0058] Then, use the bidirectional encoder representation model to perform deep feature extraction on the preprocessed data. With its powerful language understanding ability, the BERT model can convert each text unit into a high-dimensional vector. These vectors not only contain the literal meaning of the words but also incorporate the semantic information in the context, helping the model comprehensively understand the inherent meaning and potential risks of the network news content.
[0059] Secondly, after obtaining the high-dimensional feature representation, use the Transformer structure for further refinement to identify the key entities (such as organizations, products, strategies, events, etc.) in the text and the complex interaction relationships between these entities, such as subordination, chronological order, causal association, etc. This process can deeply analyze the structure and semantics of the network news data, providing detailed information support for subsequent risk assessment.
[0060] Finally, convert the identified entities and their relationships into a graph-structured data representation to form the target knowledge graph. The target knowledge graph includes the identified entities and the relationships between each entity. This data structure not only intuitively shows the constituent elements and interconnections of the network news content but also facilitates the efficient operation and in-depth analysis of subsequent risk scoring algorithms, thereby accurately evaluating the risk level of the network news and guiding the intelligent decision-making of the review process.
[0061] Optionally, in the method for identifying the risk of network news data provided in Embodiment 1 of this application, data preprocessing is performed on the network news data to be processed to obtain preprocessed data, including: using a regular expression to perform data cleaning on the network news data to be processed, performing word segmentation on the cleaned network news data, and converting the segmented data into input sequence data, where the input sequence data is in the format of the input data of the bidirectional encoder representation model; adding a first marker at the start position of the input sequence data and adding a second marker at the end position of the input sequence data to obtain the marked input sequence data; performing data padding or data truncation operations on the marked input sequence data according to a preset data length to obtain the preprocessed data.
[0062] In this Embodiment 1, data cleaning includes removing irrelevant characters, such as HTML tags, special symbols, etc. For article titles, texts, keywords, authors, and release times, regular expressions can be used to accurately remove redundant information in the network news data, improve data quality, and lay a good foundation for subsequent processing. Exemplarily, the regular expressions used for titles and texts can be as follows: cleaned_text = regex_sub(' [^a-zA-Z0-9\u4e00-\u9fa5]', '', original_text).
[0063] Then, the text is segmented into words or phrases through tokenization techniques so that the model can understand and process it more effectively. The segmented words are converted into a sequence of numbers that the model can understand (i.e., the input sequence data mentioned above, which are tokens). This process usually involves mapping each word to a unique ID to form a sequence of numbers.
[0064] Secondly, start token [CLS] (i.e., the first token mentioned above) and end token [SEP] (i.e., the second token mentioned above) are added to the beginning and end of each sequence in the input sequence data. These two tokens help the model identify the start and end of the text, which is crucial for processing and understanding the text, especially in multi-paragraph or complex text structures.
[0065] Finally, to adapt to the input length requirements of the model, overly long sequences are truncated, only retaining the most crucial parts, while shorter sequences are extended to the required length by adding padding tokens such as [PAD] to ensure that all input data has the same length, facilitating batch processing and model training.
[0066] Through the above steps, not only can the original network news data collected be standardized, but it can also be converted into an input format suitable for deep learning models (especially the BERT model). By adding key tokens, the model's understanding of the text structure and key content is optimized. Finally, preprocessed data with a unified length, clear structure, and rich information is obtained, significantly improving the efficiency and accuracy of the model in evaluating risks.
[0067] Optionally, in the method for identifying network news data risks provided in Embodiment 1 of this application, a bidirectional encoder representation model is used to extract features from the preprocessed data to obtain a high-dimensional feature representation, including: converting the preprocessed data into an embedding representation; using the self-attention mechanism in the bidirectional encoder representation model to calculate the similarity between each token in the input sequence corresponding to the embedding representation, and performing a weighted sum on the similarities to obtain the attention information of the embedding representation; using multiple encoder layers in the bidirectional encoder representation model to determine the complex context relationship of the embedding representation based on the attention information data of the embedding representation, and generating a high-dimensional feature representation.
[0068] In this Embodiment 1, it is necessary to convert the network news text data after cleaning, tokenization, and format adjustment into a set of dense digital vectors, namely the so-called embedding representation. The embedding representation captures the semantic information and context relationship of words and is a key step for deep learning models to process natural language. For each word or token, a high-dimensional vector is generated through the word embedding layer of the model, and these vectors synthesize the meaning of the word and its position information in the text.
[0069] Then, when processing the text sequence, through the self-attention mechanism in the Transformer architecture, it not only focuses on the current word but also takes into account the influence of other words in the sequence, thereby better understanding the global structure of the text. Specifically, the model calculates the similarity between all word pairs, assigns weights based on these similarities, and performs a weighted sum of the embedding vectors of each word to generate the attention representation of the word. This step strengthens the model's attention to important information in the text.
[0070] Finally, the embedding representation together with the attention information is input into multiple encoder layers. These encoder layers are stacked to further deepen the model's understanding of the text. Each encoder layer, based on the output of the previous layer, uses the attention mechanism to capture deeper context associations, thereby generating a richer high-dimensional feature representation. After this series of operations, the resulting high-dimensional feature representation can comprehensively reflect the semantic features and complex structures of the network news content, providing a data basis for subsequent named entity recognition, relation extraction, and risk assessment.
[0071] Through the above steps, the conversion from the original network news text to the high-dimensional feature representation is achieved. It not only deeply mines the semantic levels of the text but also captures and strengthens the complex associations and potential information contained in the text through the synergistic effect of the self-attention mechanism and multiple encoder layers. The finally generated high-dimensional feature representation provides a solid data basis and computational framework for the automated review and optimization of network news data of financial institutions.
[0072] Optionally, in the method for identifying network news data risks provided in Embodiment 1 of this application, converting the preprocessed data into an embedding representation includes: converting each token in the preprocessed data into a vector representation of a fixed size to obtain a token embedding vector; converting each segment in the preprocessed data into a vector representation to obtain a segment embedding vector; determining a position embedding vector according to the position of each token in the input sequence in the preprocessed data; and determining the embedding representation according to the token embedding vector, the segment embedding vector, and the position embedding vector.
[0073] In this Embodiment 1, each vocabulary or special symbol in the text information of the network news data, such as the words after word segmentation, is converted into a digital vector of a fixed length, that is, the above-mentioned token embedding vector, which can be represented as TokenEmbeddings. This vector not only contains the semantic information of the word itself but also unifies the vector dimension to ensure that the model can process it efficiently.
[0074] Then, independent vector representations, i.e., the above-mentioned segment embedding vectors, are assigned to different sentences or text segments of the text to distinguish the sources and functions of different text segments. For example, in the BERT model, different segment identifiers (Segment IDs) are used to identify different sentences or text segments of the text, which helps the model understand the interrelationships between different text segments.
[0075] Secondly, considering the impact of the position of words in a sentence on semantic understanding, the model generates a position embedding vector for each word in the text sequence to ensure that the model can identify the order and position information of words in the text, which is beneficial for the model to understand the text structure and context meaning.
[0076] Finally, the above three vectors are combined, and the final embedding representation is generated through an addition operation. The final embedding representation can be as shown in Formula One:
[0077] (InputEmbedding(t) = TokenEmbedding(t) + SegmentEmbedding(t) +
[0078] PositionEmbedding(t)) (One)
[0079] Among them, InputEmbedding(t) represents the embedding representation, TokenEmbedding(t) represents the token embedding vector, SegmentEmbedding(t) represents the segment embedding vector, and PositionEmbedding(t) represents the position embedding vector. This representation contains the semantic information of the vocabulary, the source information of the text segment, and the position information of the words, forming a comprehensive and structured vector, providing rich and organized data input for subsequent deep learning models, such as the Transformer-based bidirectional encoder model, and greatly enhancing the model's parsing ability for network news content and the accuracy of risk assessment.
[0080] Through the above steps, the preprocessed text data is converted into semantic and structural vectors that can be understood by the deep learning model, providing a solid data foundation for the bank network news review and optimization method based on the graph neural network, ensuring that the model can comprehensively capture text information from multiple dimensions and improving the accuracy of network news risk review.
[0081] Optionally, in the method for identifying network communication data risks provided in Embodiment 1 of this application, named entity recognition and relationship extraction are performed on the high-dimensional feature representation based on a transformer model to obtain multiple entities and the relationships between multiple entities, including: sequentially inputting the high-dimensional feature representation into the encoder and decoder of the transformer model to output the entity label of each token, where the entity label is used to identify different types of entities; annotating the entity corresponding to each token using a tagging algorithm based on the entity type of each token to obtain the entity annotation information of each token, and determining multiple entities based on the entity annotation information of each token; identifying entity pairs with an association relationship among the multiple entities according to the entity annotation information of each token, and using a bidirectional encoder representation model to extract the context information of each entity; inputting each entity and the context information of each entity into the transformer model to classify the relationship of the entity pair and obtain the relationships between multiple entities, where each relationship among the relationships between multiple entities includes at least one of the following: cooperation relationship, subordination relationship, and time relationship.
[0082] In this Embodiment 1, it is necessary to send the text feature vector obtained by the previous processing into the bidirectional encoder representation model, capture the context information through the self-attention mechanism of the model, and the decoder predicts the entity type represented by each token based on the information output by the encoder to generate entity labels, such as person names, organizations, locations, and dates.
[0083] Then, according to the entity labels output by the model, the BIO tagging method (Beginning, Inside, Outside, which can be simply referred to as the BIO algorithm) is used to annotate the entities in the text to identify various entities. This process is not limited to directly tagging entities, but also involves the specific identification of entity boundaries and types, ensuring the accuracy of the annotation and the integrity of entity recognition. Integrate the annotation information to identify and determine all entities appearing in the text to form an entity list.
[0084] Exemplarily, the entity types may include but are not limited to: `ORG` (organization), `PROD` (product), `LOC` (location), `STRATEGY` (strategy / planning), `EVENT` (event / conference), etc. The following is an example of BIO annotation for some entities and their relationships: 1. The First Financial Loan `B-PROD` has a rapid growth `O`. Explanation: `The First Financial Loan` as a whole is annotated as the beginning (B) of the product entity; 2. The Second Bank `B-ORG` the board of directors `O` is responsible for `O` formulating `O` the inclusive financial development plan `B-STRATEGY`. Explanation: `The Second Bank` is the beginning (B) of the organizational entity, and `The First Financial Development Plan` is the beginning (B) of the strategic plan.
[0085] Then, multiple entities, entity pairs, and their context information are fed into the Transformer structure, and the relationships between entity pairs are classified through an encoder-decoder model. The model is trained using a dataset with entity relationship annotations to learn different types of relationships, such as colleagues, location associations, etc. Relationship extraction not only includes direct entity markers but also further analyzes the types of relationships existing between these entities, such as cooperation relationships, subordination relationships, time relationships, etc.
[0086] Through the above steps, the entities and relationships between entities in the network communication data text can be deeply analyzed, providing the core logic and data preparation for implementing the adaptive optimization method of bank network communication audit based on graph neural networks, and greatly improving the intelligence and accuracy of the audit process.
[0087] Optionally, in the method for identifying network communication data risks provided in Embodiment 1 of the present application, a graph structure is used to represent multiple entities and the relationships between multiple entities, obtaining a target knowledge graph, including: creating corresponding nodes in the graph according to each entity in the multiple entities; creating corresponding edges in the graph according to the entity pairs with existing relationships in the entity pairs; creating an initial knowledge graph according to the nodes and edges, and constructing an adjacency matrix according to the initial knowledge graph; optimizing the nodes and edges included in the initial knowledge graph, updating the adjacency matrix, determining the target knowledge graph according to the updated adjacency matrix, and storing the target knowledge graph in a preset storage space.
[0088] In this Embodiment 1, each entity (such as an organization, a product, a location, etc.) identified in the text analysis stage is represented as an independent node in the knowledge graph. These nodes are the basic components of the graph and represent the entities themselves. The creation of nodes can be expressed as: (Node i CreateNode=(Entity i ))), where Entity i represents the i-th entity. By identifying the relationships between entities, such as cooperation, subordination, or temporal associations, the nodes connecting the relevant entities in the graph are formed into edges. The existence of edges not only represents the connection between entities but also carries information about the relationship type. The creation of edges can be expressed as: (Edge ij CreateEdge=.Entity i ,Entity j ,Relation ij / ), where Entity i represents the i-th entity, Entity j represents the j-th entity, representing the relationship between the i-th entity and the j-th entity.
[0089] Then, all the nodes and edges are integrated together to construct an initial graphical structure, i.e., the initial knowledge graph. The integration of the graph can be represented by the data structure of the graph. For example, an adjacency matrix or an adjacency list. The adjacency matrix is used to record the direct and indirect connection relationships between the nodes in the graph, facilitating subsequent processing and analysis by the graph neural network. After the graph construction is completed, the graph is adjusted and improved by optimizing the nodes and edges, such as deleting isolated nodes, merging duplicate edges, removing redundant information, or adding missing links. This process will update the adjacency matrix in real time to ensure that it accurately reflects the optimized graph structure.
[0090] Finally, the final target knowledge graph is formed based on the optimization results. The target knowledge graph not only contains entities and their explicit relationships, but also contains implicit connections hidden between the context information. By storing it in a preset space, it provides convenience for subsequent querying, analysis, and model training, and at the same time ensures the persistence and accessibility of the knowledge graph, providing a strong data support and analysis framework for the adaptive review optimization of bank network information.
[0091] Optionally, in the method for identifying the risk of network information provided in Embodiment 1 of this application, based on the gradient descent algorithm, the target evaluation model optimized for the Nth time is optimized according to the review results of historical risk data to obtain the target evaluation model optimized for the (N + 1)th time, including: receiving the review results of the target object for the historical risk data, and annotating the review results to obtain the annotation information of the review results; performing data preprocessing on the review results and the annotation information of the review results to obtain the preprocessed review information; based on the gradient descent algorithm, using the preprocessed review information to optimize the model parameters of the target evaluation model optimized for the Nth time to obtain the target evaluation model optimized for the (N + 1)th time.
[0092] In this Embodiment 1, the review judgments of the historical risk data from the target object, that is, professional reviewers, are received. These judgments include the evaluation of the compliance, security, and content integrity of the information. Then, these review results are carefully annotated to create annotation information. This step is to convert the reviewers' decisions into machine-readable labels, such as marking the information as categories like "high risk", "low risk", or "compliant" for subsequent use in training the model.
[0093] Next, data preprocessing is performed on the annotated review results and their corresponding labels, specifically including data cleaning, text conversion to numerical labels, format standardization, and feature engineering, which not only ensures the data quality but also improves the efficiency of model training. The preprocessed review information not only removes the noise and irrelevant information in the original data but also makes the data more suitable for the input requirements of the deep learning model through conversion and encoding.
[0094] Finally, based on the gradient descent algorithm, using the above-mentioned preprocessed review information, update the parameters of the target evaluation model that has completed the Nth iteration of optimization to obtain the optimized model after the (N + 1)th iteration. Adjust the learning rate and other hyperparameters according to the performance of the model. For example, if the performance of the model on new data deteriorates, it may be necessary to reduce the learning rate to avoid overfitting. The process of model update can be expressed as: where (w new ) and (w old ) are the weights before and after the update respectively, (α) is the learning rate, is the gradient of the loss function (L) with respect to (w old ). The cross-entropy function can be used as the loss function.
[0095] Furthermore, the iterative optimization operation of the target evaluation model can also be triggered according to time or model accuracy. Repeat the above steps, continuously iterate the model to adapt to new data and environmental changes. Use techniques such as cross-validation to evaluate the generalization ability of the model, and adjust the model structure accordingly, so that the model gradually learns to identify risk features from the preprocessed review information, thereby improving the accuracy and robustness of its risk assessment. Through this closed-loop optimization process, the model can continuously learn from new feedback, thereby improving the accuracy and adaptability of its risk assessment of network communication content.
[0096] Through the above steps, a closed-loop learning and optimization process is formed, aiming to continuously improve the review ability of the target evaluation model for historical risk data to adapt to the continuously changing risk types and compliance requirements in the financial environment, and ensure the safe review and compliant dissemination of bank network communication content.
[0097] Optionally, in this Embodiment 1, Figure 3 is a schematic diagram of an optional method for identifying network communication data risks provided in this Embodiment 1. Figure 3 It presents the entire workflow starting from data preprocessing, involving the cleaning and formatting of the original network communication data, preparing a high-quality and standardized dataset for model input. Then, the BERT model is applied to process the pre-cleaned data to generate high-dimensional feature representations, which contain rich semantic information of the text and provide a solid foundation for subsequent in-depth analysis. Secondly, named entity recognition (NER) and relation extraction techniques extract entities and their interrelationships in the high-dimensional features, constructing a network structure between entities, which provides strong support for understanding the micro details of network communication content. Finally, the knowledge graph construction stage visualizes these entities and relationships to form a graph structure, providing a structured data basis for further using graph neural networks for risk assessment. Immediately afterwards, Figure 3The paper shows how a bidirectional graph neural network (Bi-GNN) analyzes the constructed knowledge graph and predicts the risk level of network information data. This process utilizes the bidirectional propagation of complex relationships between entities to improve the accuracy of risk assessment. Then, the adaptive optimization strategy dynamically adjusts the audit process based on the prediction results of Bi-GNN to ensure the efficiency and security of the audit. Secondly, the feedback loop mechanism is used in Figure 3 It is emphasized that by collecting feedback data from manual review and continuously optimizing model parameters, the model can continuously learn and adapt to new risk characteristics. Finally, this closed-loop optimization process ensures the long-term effectiveness and adaptability of the model, and through continuous iterative improvement, maintains the advancement and flexibility of the bank's online information review mechanism to cope with the ever-changing challenges of financial information security.
[0098] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0099] In summary, the method for identifying network information data risks provided in the embodiment of the present application calculates the historical knowledge graph corresponding to the historical network information data through the target evaluation model optimized for the Nth time to obtain a first risk score, and determines the historical risk data based on the first risk score, wherein N is a non-negative integer; based on the gradient descent algorithm, the target evaluation model optimized for the Nth time is optimized according to the audit results of the historical risk data to obtain the target evaluation model optimized for the N+1th time; a bidirectional encoder representation model is used to process the current network information data to be processed to construct a target knowledge graph; the target knowledge graph is input into the target evaluation model optimized for the N+1th time, the risk score of the network information data to be processed is calculated, and a risk assessment result is generated, which solves the problem that when traditional text analysis and rule-based methods are used in related technologies to identify network information data risks, the online learning and instant feedback mechanism is missing, which makes the model unable to self-optimize and leads to inaccurate risk identification results.
[0100] Calculating the historical knowledge graph corresponding to the historical network news data through the target evaluation model optimized for the Nth time can accurately quantify the risk level of historical content, achieving the technical effect of refined risk identification. At the same time, based on the gradient descent algorithm, the (N + 1)th iterative optimization is performed according to the audit results of historical risk data, enabling the fine-tuning of model parameters, realizing the self-correction and efficiency improvement of the model, and further achieving the technical effect of enhancing the model's adaptability and prediction stability; using a bidirectional encoder representation model to process the current network news data to be processed and constructing a target knowledge graph can deeply mine and structure the entities and relationships in new information, achieving the technical effect of intelligent parsing and efficient information integration. Inputting this knowledge graph into the target evaluation model optimized for the (N + 1)th time to calculate the risk score of the network news data to be processed and generate a risk assessment result can real-time evaluate the content risk, realizing intelligent auditing and instant response, and further achieving the technical effect of improving the auditing efficiency and ensuring the security of information dissemination.
[0101] Embodiment 2
[0102] The embodiment of the present application further provides a device for identifying the risk of network news data. It should be noted that the device for identifying the risk of network news data in the embodiment of the present application can be used to execute the method for identifying the risk of network news data provided by the embodiment of the present application. The following introduces the device for identifying the risk of network news data provided by the embodiment of the present application.
[0103] According to the embodiment of the present application, there is also provided a device for implementing the above method for identifying the risk of network news data. Figure 4 is a schematic diagram of the device for identifying the risk of network news data provided according to Embodiment 2 of the present application, as Figure 4 shown, the device includes:
[0104] Specifically, a calculation unit 401 is configured to calculate the historical knowledge graph corresponding to the historical network news data through the target evaluation model optimized for the Nth time to obtain a first risk score, and determine historical risk data based on the first risk score, where N is a non-negative integer.
[0105] An optimization unit 402 is configured to optimize the target evaluation model optimized for the Nth time based on the gradient descent algorithm according to the audit results of historical risk data to obtain the target evaluation model optimized for the (N + 1)th time.
[0106] A construction unit 403 is configured to process the current network news data to be processed by using a bidirectional encoder representation model to construct a target knowledge graph.
[0107] A generation unit 404 is configured to input the target knowledge graph into the target evaluation model optimized for the (N + 1)th time, calculate the risk score of the network news data to be processed, and generate a risk assessment result.
[0108] The risk identification device for network communication data provided by the embodiment of the present application calculates, through the target evaluation model optimized for the Nth time by the calculation unit 401, the historical knowledge graph corresponding to the historical network communication data, obtains the first risk score, and determines the historical risk data according to the first risk score, where N is a non-negative integer; the optimization unit 402 optimizes the target evaluation model optimized for the Nth time according to the review result of the historical risk data based on the gradient descent algorithm to obtain the target evaluation model optimized for the (N + 1)th time; the construction unit 403 processes the currently to-be-processed network communication data by using a bidirectional encoder representation model to construct a target knowledge graph; the generation unit 404 inputs the target knowledge graph into the target evaluation model optimized for the (N + 1)th time, calculates the risk score of the to-be-processed network communication data, and generates a risk assessment result, solving the problem that in the related art, when identifying the risk of network communication data by using traditional text analysis and rule-based methods, the online learning and instant feedback mechanism are missing, so that the model cannot optimize itself, resulting in inaccurate risk identification results.
[0109] Calculating the historical knowledge graph corresponding to the historical network communication data through the target evaluation model optimized for the Nth time can accurately quantify the risk level of the historical content, achieving the technical effect of refined risk identification. At the same time, through the (N + 1)th iterative optimization based on the gradient descent algorithm according to the review result of the historical risk data, the model parameters are finely tuned, realizing the self-correction and efficiency improvement of the model, and further achieving the technical effect of enhancing the adaptability and prediction stability of the model; using a bidirectional encoder representation model to process the currently to-be-processed network communication data to construct a target knowledge graph can deeply excavate and structure the entities and relationships in the new information, achieving the technical effect of intelligent parsing and efficient information integration. Inputting this knowledge graph into the target evaluation model optimized for the (N + 1)th time, calculating the risk score of the to-be-processed network communication data, and generating a risk assessment result can real-time evaluate the content risk, realizing intelligent review and instant response, and further achieving the technical effect of improving the review efficiency and ensuring the security of information dissemination.
[0110] Optionally, in the risk identification device for network communication data provided in the second embodiment of the present application, the above-mentioned construction unit 403 includes: a first processing subunit, configured to perform data preprocessing on the to-be-processed network communication data to obtain preprocessed data; an extraction subunit, configured to extract features from the preprocessed data by using a bidirectional encoder representation model to obtain a high-dimensional feature representation; an identification subunit, configured to perform named entity recognition and relationship extraction on the high-dimensional feature representation based on a transformer model to obtain multiple entities and the relationships between multiple entities; a characterization subunit, configured to characterize the multiple entities and the relationships between multiple entities by using a graph structure to obtain a target knowledge graph.
[0111] Optionally, in the network communication data risk identification device provided in the second embodiment of the present application, the above-mentioned processing subunit includes: a cleaning module, configured to perform data cleaning on the network communication data to be processed by using a regular expression, perform word segmentation on the cleaned network communication data, and convert the segmented data into input sequence data, where the input sequence data is in the format of the input data of the bidirectional encoder representation model; an adding module, configured to add a first marker at the start position of the input sequence data and add a second marker at the end position of the input sequence data to obtain the marked input sequence data; a padding module, configured to perform data padding operation or data truncation operation on the marked input sequence data according to a preset data length to obtain preprocessed data.
[0112] Optionally, in the network communication data risk identification device provided in the second embodiment of the present application, the above-mentioned extraction subunit includes: a conversion module, configured to convert the preprocessed data into an embedded representation; a calculation module, configured to calculate the similarity between each marker in the input sequence corresponding to the embedded representation by using the self-attention mechanism in the bidirectional encoder representation model, and perform weighted summation on the similarity to obtain the attention information of the embedded representation; a generation module, configured to use multiple encoder layers in the bidirectional encoder representation model to determine the complex context relationship of the embedded representation according to the attention information data of the embedded representation and generate a high-dimensional feature representation.
[0113] Optionally, in the network communication data risk identification device provided in the second embodiment of the present application, the above-mentioned conversion module includes: a first conversion sub-module, configured to convert each marker in the preprocessed data into a vector representation of a fixed size to obtain a marker embedding vector; a second conversion sub-module, configured to convert each segment in the preprocessed data into a vector representation to obtain a segment embedding vector; a first determination sub-module, configured to determine a position embedding vector according to the position of each marker in the input sequence in the preprocessed data; a second determination sub-module, configured to determine the embedded representation according to the marker embedding vector, the segment embedding vector, and the position embedding vector.
[0114] Optionally, in the network communication data risk identification device provided in the second embodiment of this application, the above-mentioned identification subunit includes: an output module, configured to sequentially input the high-dimensional feature representation into the encoder and decoder of the converter model, and output the entity label of each token, where the entity label is used to identify different types of entities; a labeling module, configured to label the entity corresponding to each token by using a labeling algorithm based on the entity type of each token, obtain the entity labeling information of each token, and determine multiple entities according to the entity labeling information of each token; an identification module, configured to identify the entity pairs with an association relationship among the multiple entities according to the entity labeling information of each token, and extract the context information of each entity by using a bidirectional encoder representation model; a classification module, configured to input each entity and the context information of each entity into the converter model, classify the relationship of the entity pair, and obtain the relationship among the multiple entities, where each relationship among the multiple entities includes at least one of the following: a cooperation relationship, a subordination relationship, and a time relationship.
[0115] Optionally, in the network communication data risk identification device provided in the second embodiment of this application, the above-mentioned characterization subunit includes: a first creation module, configured to create a corresponding node in the graph according to each entity among the multiple entities; a second creation module, configured to create a corresponding edge in the graph according to the entity pairs with a relationship in the entity pairs; a third creation module, configured to create an initial knowledge graph according to the nodes and edges, and construct an adjacency matrix according to the initial knowledge graph; a storage module, configured to optimize the nodes and edges included in the initial knowledge graph, update the adjacency matrix, determine the target knowledge graph according to the updated adjacency matrix, and store the target knowledge graph in a preset storage space.
[0116] Optionally, in the network communication data risk identification device provided in the second embodiment of this application, the above-mentioned optimization unit 402 includes: a labeling subunit, configured to receive the review result of the target object for the historical risk data, and label the review result to obtain the labeled information of the review result; a second processing subunit, configured to perform data preprocessing on the review result and the labeled information of the review result to obtain the preprocessed review information; an optimization subunit, configured to optimize the model parameters of the target evaluation model optimized for the Nth time by using the preprocessed review information based on the gradient descent algorithm to obtain the target evaluation model optimized for the (N + 1)th time.
[0117] It should be noted here that the above calculation unit 401, optimization unit 402, construction unit 403, and generation unit 404 correspond to steps S201 to S204 in Embodiment 1. The instances and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules or units can be hardware components or software components stored in a memory (for example, memory 104) and processed by one or more processors (for example, processors 102a, 102 Network communication data risk identification,..., 102n). The above modules can also be part of a device and can run in the computer terminal 10 provided in Embodiment 1.
[0118] Embodiment 3
[0119] An embodiment of the present application may provide an electronic device. Figure 5 It is a schematic diagram of an electronic device for identifying network communication data risks according to Embodiment 3 of the present application. As Figure 5 shown, the electronic device may include: one or more ( Figure 5 only one is shown in the figure) processors 502, a memory 504, a storage controller, and a peripheral interface. Among them, the peripheral interface is connected to a radio frequency module, an audio module, and a display.
[0120] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, the above methods are implemented. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely provided with respect to the processor, and these remote memories can be connected to the terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0121] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: calculate the first risk score for the historical knowledge graph corresponding to the historical network data through the target evaluation model optimized for the Nth time, and determine the historical risk data based on the first risk score, where N is a non-negative integer; optimize the target evaluation model optimized for the Nth time based on the audit results of the historical risk data according to the gradient descent algorithm to obtain the target evaluation model optimized for the (N + 1)th time; use the bidirectional encoder representation model to process the currently to-be-processed network data to construct a target knowledge graph; input the target knowledge graph into the target evaluation model optimized for the (N + 1)th time, calculate the risk score of the to-be-processed network data, and generate a risk assessment result.
[0122] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: use the bidirectional encoder representation model to process the currently to-be-processed network data to construct a target knowledge graph, including: performing data preprocessing on the to-be-processed network data to obtain preprocessed data; using the bidirectional encoder representation model to extract features from the preprocessed data to obtain a high-dimensional feature representation; performing named entity recognition and relationship extraction on the high-dimensional feature representation based on the transformer model to obtain multiple entities and the relationships between multiple entities; using a graph structure to represent the multiple entities and the relationships between multiple entities to obtain a target knowledge graph.
[0123] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: perform data preprocessing on the to-be-processed network data to obtain preprocessed data, including: using a regular expression to clean the to-be-processed network data, segmenting the cleaned network data, and converting the segmented data into input sequence data, where the input sequence data is in the format of the input data of the bidirectional encoder representation model; adding a first marker at the start position of the input sequence data and adding a second marker at the end position of the input sequence data to obtain the marked input sequence data; performing a data padding operation or a data truncation operation on the marked input sequence data according to a preset data length to obtain the preprocessed data.
[0124] The processor can call the information and application programs stored in the memory through a transmission device to perform the following steps: extracting features from the preprocessed data using a bidirectional encoder representation model to obtain a high-dimensional feature representation, including: converting the preprocessed data into an embedding representation; calculating the similarity between each token in the input sequence corresponding to the embedding representation using the self-attention mechanism in the bidirectional encoder representation model, and performing a weighted sum on the similarity to obtain the attention information of the embedding representation; using multiple encoder layers in the bidirectional encoder representation model to determine the complex context relationship of the embedding representation based on the attention information data of the embedding representation, and generating a high-dimensional feature representation.
[0125] The processor can call the information and application programs stored in the memory through a transmission device to perform the following steps: converting the preprocessed data into an embedding representation, including: converting each token in the preprocessed data into a fixed-size vector representation to obtain a token embedding vector; converting each segment in the preprocessed data into a vector representation to obtain a segment embedding vector; determining a position embedding vector based on the position of each token in the input sequence in the preprocessed data; and determining the embedding representation based on the token embedding vector, the segment embedding vector, and the position embedding vector.
[0126] The processor can call the information and application programs stored in the memory through a transmission device to perform the following steps: performing named entity recognition and relationship extraction on the high-dimensional feature representation based on a transformer model to obtain multiple entities and the relationships between multiple entities, including: sequentially inputting the high-dimensional feature representation into the encoder and decoder of the transformer model to output the entity label of each token, where the entity label is used to identify different types of entities; annotating the entity corresponding to each token using a token algorithm based on the entity type of each token to obtain the entity annotation information of each token, and determining multiple entities based on the entity annotation information of each token; identifying entity pairs with an association relationship among multiple entities based on the entity annotation information of each token, and extracting the context information of each entity using a bidirectional encoder representation model; inputting each entity and the context information of each entity into the transformer model to classify the relationship of the entity pair and obtain the relationships between multiple entities, where each relationship in the relationships between multiple entities includes at least one of the following: cooperation relationship, subordination relationship, time relationship.
[0127] The processor can call the information and application programs stored in the memory through a transmission device to execute the following steps: representing the relationships between multiple entities and multiple entities using a graph structure to obtain a target knowledge graph, including: creating corresponding nodes in the graph based on each entity in the multiple entities; creating corresponding edges in the graph based on entity pairs with existing relationships; creating an initial knowledge graph based on the nodes and edges, and constructing an adjacency matrix based on the initial knowledge graph; optimizing the nodes and edges included in the initial knowledge graph, updating the adjacency matrix, determining the target knowledge graph based on the updated adjacency matrix, and storing the target knowledge graph in a preset storage space.
[0128] The processor can call the information and application programs stored in the memory through a transmission device to execute the following steps: optimizing the target evaluation model optimized for the Nth time based on the audit results of historical risk data according to the gradient descent algorithm to obtain the target evaluation model optimized for the (N + 1)th time, including: receiving the audit results of the target object for the historical risk data and annotating the audit results to obtain the annotation information of the audit results; performing data preprocessing on the audit results and the annotation information of the audit results to obtain the preprocessed audit information; optimizing the model parameters of the target evaluation model optimized for the Nth time using the preprocessed audit information according to the gradient descent algorithm to obtain the target evaluation model optimized for the (N + 1)th time.
[0129] Using the embodiments of the present application, a method for identifying risks in network communication data is provided. Through E, it further solves the technical problem that when using traditional text analysis and rule-based methods to identify risks in network communication data, the lack of an online learning and instant feedback mechanism makes the model unable to self-optimize, resulting in inaccurate risk identification results.
[0130] Calculating the historical knowledge graph corresponding to the historical network communication data through the target evaluation model optimized for the Nth time can accurately quantify the risk level of the historical content, achieving the technical effect of refined risk identification. At the same time, through the (N + 1)th iterative optimization based on the audit results of historical risk data according to the gradient descent algorithm, the model parameters can be finely tuned, realizing the self-correction and efficiency improvement of the model, and further achieving the technical effect of enhancing the adaptability and prediction stability of the model; using a bidirectional encoder representation model to process the current network communication data to be processed and constructing a target knowledge graph can deeply mine and structure the entities and relationships in the new information, achieving the technical effect of intelligent parsing and efficient information integration. Inputting this knowledge graph into the target evaluation model optimized for the (N + 1)th time to calculate the risk score of the network communication data to be processed and generate a risk assessment result can real-time evaluate the content risk, realizing intelligent auditing and instant response, and further achieving the technical effect of improving the auditing efficiency and ensuring the security of information dissemination.
[0131] Those of ordinary skill in the art can understand that Figure 5 the structure shown is only illustrative, and the electronic device can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and mobile Internet devices (Mobile Internet Devices, MID), PAD and other terminal devices. Figure 5 It does not limit the structure of the above electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, a display device, etc.) than those shown Figure 5 in it, or have a different configuration from that shown Figure 5 in it.
[0132] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc, etc.
[0133] Embodiment 4
[0134] The embodiment of the present application further provides a storage medium. Optionally, in this embodiment, the above storage medium can be used to save the program code executed by the method for identifying the risk of network communication data provided in the above Embodiment 1.
[0135] Optionally, in this embodiment, the above storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0136] The present application also provides a computer program product, which is suitable for executing a program for the steps of the method for identifying the risk of network communication data when executed on a data processing device.
[0137] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0138] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0139] In several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0140] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0141] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0142] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0143] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A method for identifying risks in network communication data, characterized in that, Including: Calculating a first risk score for a historical knowledge graph corresponding to historical network news data by using a target evaluation model optimized for the Nth time, and determining historical risk data based on the first risk score, where N is a non - negative integer; Optimizing the target evaluation model optimized for the Nth time based on the audit result of the historical risk data according to the gradient descent algorithm to obtain a target evaluation model optimized for the (N + 1)th time; Processing the currently to - be - processed network news data by using a bidirectional encoder representation model to construct a target knowledge graph; Inputting the target knowledge graph into the target evaluation model optimized for the (N + 1)th time, calculating the risk score of the to - be - processed network news data, and generating a risk assessment result.
2. The method according to claim 1, wherein Processing the currently to - be - processed network news data by using a bidirectional encoder representation model to construct a target knowledge graph, including: Performing data pre - processing on the to - be - processed network news data to obtain pre - processed data; Using the bidirectional encoder representation model to extract features from the pre - processed data to obtain a high - dimensional feature representation; Based on a transformer model, performing named entity recognition and relationship extraction on the high - dimensional feature representation to obtain multiple entities and the relationships between the multiple entities; Using a graph structure to represent the multiple entities and the relationships between the multiple entities to obtain the target knowledge graph.
3. The method according to claim 2, wherein Performing data pre - processing on the to - be - processed network news data to obtain pre - processed data, including: Using a regular expression to clean the to - be - processed network news data, segmenting the cleaned network news data, and converting the segmented data into input sequence data, where the input sequence data is in the format of the input data of the bidirectional encoder representation model; Adding a first marker at the start position of the input sequence data and adding a second marker at the end position of the input sequence data to obtain marked input sequence data; Performing a data padding operation or a data truncation operation on the marked input sequence data according to a preset data length to obtain the pre - processed data.
4. The method according to claim 2, wherein Using the bidirectional encoder representation model to extract features from the pre - processed data to obtain a high - dimensional feature representation, including: Converting the pre - processed data into an embedding representation; Using the self - attention mechanism in the bidirectional encoder representation model to calculate the similarity between each marker in the input sequence corresponding to the embedding representation, and performing a weighted sum on the similarity to obtain the attention information of the embedding representation; Using multiple encoder layers in the bidirectional encoder representation model to determine the complex context relationship of the embedding representation based on the attention information data of the embedding representation, and generating the high - dimensional feature representation.
5. The method according to claim 4, wherein Converting the pre - processed data into an embedding representation, including: Converting each marker in the pre - processed data into a fixed - size vector representation to obtain a marker embedding vector; Converting each segment in the pre - processed data into a vector representation to obtain a segment embedding vector; Determining a position embedding vector according to the position of each marker in the input sequence of the pre - processed data; Determine the embedding representation based on the marked embedding vector, the segment embedding vector, and the position embedding vector.
6. The method according to claim 2, wherein Perform named entity recognition and relationship extraction on the high-dimensional feature representation based on a transformer model to obtain multiple entities and the relationships between the multiple entities, including: Input the high-dimensional feature representation into the encoder and decoder of the transformer model in sequence, and output the entity label of each token, where the entity label is used to identify different types of entities; Annotate the entity corresponding to each token using a token algorithm based on the entity type of each token to obtain the entity annotation information of each token, and determine the multiple entities based on the entity annotation information of each token; Identify entity pairs with an association relationship among the multiple entities based on the entity annotation information of each token, and extract the context information of each entity using the bidirectional encoder representation model; Input each entity and the context information of each entity into the transformer model to classify the relationship of the entity pair and obtain the relationships between the multiple entities, where each relationship among the relationships between the multiple entities includes at least one of the following: cooperation relationship, subordination relationship, time relationship.
7. The method according to claim 2, wherein Characterize the multiple entities and the relationships between the multiple entities using a graph structure to obtain the target knowledge graph, including: Create corresponding nodes in the graph according to each entity among the multiple entities; Create corresponding edges in the graph according to the entity pairs with relationships in the entity pairs; Create an initial knowledge graph according to the nodes and the edges, and construct an adjacency matrix according to the initial knowledge graph; Optimize the nodes and edges included in the initial knowledge graph, update the adjacency matrix, determine the target knowledge graph according to the updated adjacency matrix, and store the target knowledge graph in a preset storage space.
8. The method according to claim 1, wherein Optimize the target evaluation model after the Nth optimization based on the audit results of the historical risk data according to the gradient descent algorithm to obtain the target evaluation model after the (N + 1)th optimization, including: Receive the audit results of the target object for the historical risk data, and annotate the audit results to obtain the annotation information of the audit results; Perform data preprocessing on the audit results and the annotation information of the audit results to obtain the preprocessed audit information; Optimize the model parameters of the target evaluation model after the Nth optimization using the preprocessed audit information according to the gradient descent algorithm to obtain the target evaluation model after the (N + 1)th optimization.
9. An identification device for network communication data risks, characterized in that Include: A calculation unit for calculating the historical knowledge graph corresponding to the historical network news data through the target evaluation model after the Nth optimization to obtain a first risk score, and determining the historical risk data according to the first risk score; An optimization unit for optimizing the target evaluation model after the Nth optimization based on the audit results of the historical risk data according to the gradient descent algorithm to obtain the target evaluation model after the (N + 1)th optimization; A construction unit for processing the currently to-be-processed network news data using the bidirectional encoder representation model to construct a target knowledge graph; A generating unit is configured to input the target knowledge graph into the target evaluation model optimized for the (N + 1)-th time, calculate a risk score of the network communication data to be processed, and generate a risk assessment result.
10. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by a processor, the steps of the method for identifying the risk of network communication data according to any one of claims 1 to 8 are implemented.