A knowledge graph-based web content security processing method

By combining multimodal data fusion and adaptive acquisition technology based on knowledge graphs, we build a web content security processing system, which addresses the shortcomings of traditional methods in identifying and processing complex web content security threats, achieves efficient and accurate security analysis and monitoring, provides an intuitive visual interface, and ensures continuous security and flexibility.

CN119670070BActive Publication Date: 2025-10-14CHINA DATACOM CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411712742.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-10-14
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

Traditional web content security processing methods are unable to fully cope with diverse web content data, inaccurately identify security threats, and lack knowledge integration and utilization, resulting in insufficient depth and accuracy in security analysis, making it difficult to effectively respond to complex network security threats.

Method used

It adopts a knowledge graph-based approach, through multimodal data fusion, adaptive collection, entity extraction and graph construction, combined with deep learning and visualization technology, to achieve comprehensive, accurate, efficient and secure processing of web page content, with dynamic update and semantic expansion capabilities.

Benefits of technology

It achieves comprehensive monitoring and analysis of web page content, improves the accuracy of security threat identification, provides an intuitive visual interface, facilitates rapid response, reduces data processing costs, and ensures the continued effectiveness of security processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119670070B_ABST
    Figure CN119670070B_ABST
Patent Text Reader

Abstract

The application discloses a kind of knowledge graph-based webpage content security processing method, it is related to webpage content security technical field, the method includes the following specific steps: data acquisition: from network service provider, enterprise or organization internal network and website server collect the internet content data related to webpage of text, image and video form, the present application realizes the comprehensive monitoring and analysis to webpage content security by constructing knowledge graph-based webpage content security processing system, by using the inference ability of knowledge graph, including representation learning, correlation analysis, event traceability, behavior prediction etc., potential security threats, abnormal patterns or associated relationships in webpage content can be deeply mined, not only improve the accuracy of security threat identification, but also provide intuitive, easy-to-understand visual interface for security personnel, so that they can quickly respond and take appropriate security measures.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of webpage content security, in particular to a webpage content security processing method based on a knowledge graph. BACKGROUND

[0002] Under the current Internet environment, webpage content is growing explosively, and its data types are becoming increasingly diverse, covering text, images, videos and other complex forms. As an important carrier of information dissemination, the security of webpage content is directly related to the health of the Internet ecosystem and the safety of user information. With the continuous progress of network technology, webpage content security processing has become an important issue in the field of network security, aiming to identify and prevent security threats such as malicious content, sensitive information leakage, and network attacks through technical means.

[0003] Traditional webpage content security processing methods have many limitations. First, data processing is not comprehensive, making it difficult to deal with massive and diverse webpage content data, resulting in missed identification of security threats. Second, security threat identification is not accurate, as traditional technical means often rely on fixed rule libraries or feature matching, making it difficult to effectively deal with changing network security threats. In addition, traditional methods lack the ability to integrate and utilize knowledge, failing to fully utilize the relevance and contextual information between webpage content, resulting in insufficient depth and accuracy of security analysis. These defects limit the effectiveness of traditional methods in dealing with complex and diverse webpage content security threats.

[0004] To address the above problems, it is necessary to optimize existing webpage content security processing methods by comprehensively collecting, processing and analyzing webpage content-related data, utilizing the representation and reasoning capabilities of knowledge graphs to effectively protect webpage content security. Therefore, it is of great significance to develop a webpage content security processing method based on a knowledge graph that can comprehensively realize the above features. SUMMARY

[0005] The present application aims to overcome the shortcomings of the prior art and provides a webpage content security processing method based on a knowledge graph. This method can achieve comprehensive, accurate and efficient security processing of webpage content by introducing adaptive acquisition strategies for multi-modal data fusion, ontology modeling, knowledge extraction, knowledge fusion, graph construction and monitoring analysis. In addition, the present application has the ability of dynamic updating and semantic expansion, which can timely adapt to new network security threats and situations, ensuring the continuous effectiveness and accuracy of webpage content security processing.

[0006] To solve the above technical problems, the present application provides the following technical solution: a webpage content security processing method based on a knowledge graph, comprising the following specific steps:

[0007] Data collection: Collecting web-related internet content data in the form of text, images, and videos from network service providers, enterprise or organization internal networks and website servers. According to the characteristics of data from different sources and the requirements of web content security analysis, adaptive collection frequency and method are developed. Real-time fusion processing of different modal data is carried out during collection using multi-modal data fusion technology, so that the collected data can meet the requirements of subsequent analysis;

[0008] Ontology modeling: Clearly define concepts related to web content security and accurately define their categories and characteristics. Determine the parent-child relationship between concepts to build a hierarchical concept system. Assign specific attribute properties to each concept, define its connotation, boundary, value range, and unique constraint conditions. Define various semantic relationships and their meanings and judgment criteria. Develop a unified naming specification, set axioms and rules, and specify trigger conditions and execution logic. Establish a semantic extension and dynamic update mechanism based on new situations to automatically extend the semantics of related concepts and review and update the ontology model at preset intervals or under specific trigger conditions;

[0009] Knowledge extraction: Use natural language processing techniques to preprocess different types of data to ensure consistency in format and feature representation. Build a joint extraction framework that integrates multi-source heterogeneous knowledge. Use the multi-head attention mechanism in deep learning to focus on key information from different data sources and types simultaneously. This allows for accurate extraction of entity, relationship, and attribute information, providing a high-quality knowledge base for subsequent knowledge fusion and graph construction;

[0010] Knowledge fusion: Use similarity functions and vector space models to calculate the similarity between entities in extracted knowledge. Align and merge entities that meet the similarity threshold as the same entity. When there are multiple different concepts or attributes corresponding to the aligned entities, use specific algorithms and logic to eliminate ambiguity by analyzing semantic relationships and context information;

[0011] Graph construction: Choose a graph database and accurately import entities and relationships between entities into the graph database after knowledge fusion according to their format requirements through specific input interfaces and data conversion processes. Construct a knowledge graph that forms a semantic network containing entities and relationships between entities. Clearly define the storage and representation of each entity and relationship in the knowledge graph;

[0012] Monitoring and analysis: Based on the constructed knowledge graph, set specific analysis algorithms and logic to deeply mine and analyze the entities and relationships in the knowledge graph. Explore potential security threats, abnormal patterns, or associated relationships;

[0013] Results display and response: using visualization technology and tools to design layout, display and interactive functions, display knowledge graph relationship and monitoring analysis results, so that safety personnel can take appropriate measures according to the results.

[0014] Further, in the data collection step, according to the characteristics of data from different sources and the safety analysis requirements of web content, adaptive collection frequency and method are formulated. For text data, the update frequency is , the importance level coefficient is , the preset basic collection frequency is , the adjustment coefficient is , and the actual collection frequency is calculated as follows: For image or video data, the data size is , the resolution is , the preset maximum sampling ratio is , the reference data size is , the reference resolution is , the adjustment coefficient is , and the actual sampling ratio is calculated as follows: .

[0015] Further, in the data collection step, multi-modal data fusion technology is used to process different modal data in real time during collection. Specifically, the collected text data vector is represented as , is the feature dimension of text data, the image data vector after feature extraction is represented as , is the feature dimension of image data, and the video data vector after processing is represented as , is the feature dimension of video data. The fusion weight coefficient , , γ and satisfy , the multi-modal data fusion vector is calculated as follows: wherein , , represent the modulus of text data vector , image data vector , and video data vector , respectively.

[0016] Further, in the ontology modeling step, the parent-child relationship between concepts is determined to build a clear hierarchical concept system. Specifically, let the two concepts to be determined be concept and concept , is the feature similarity index between concept , while there exists a set of concept relationship impact factors , represents the value of the th impact factor, is the number of impact factors, then the concept has a hierarchical relationship tendency value relative to concept , and the calculation formula is: , is the weight coefficient of the th impact factor, and satisfies , represents the hierarchical relationship tendency value of concept relative to concept , and its value range is between [0, 1], when the value of is close to 1, it means that concept is more inclined to be a subclass of concept , when the value is close to 0, it means that concept is not a subclass of concept , is the feature similarity index between concept and concept .

[0017] Further, in the ontology modeling step, is the feature similarity index between concept and concept , and the calculation formula is: , wherein, when the value of is close to 1, it means that the two concepts are highly similar in features, and when the value is close to 0, it means that the feature similarity is very low, is the number of feature dimensions, and are the feature values of concept and concept in the th feature dimension, respectively.

[0018] Further, in the ontology modeling step, the related concepts are automatically expanded in semantics according to new situations, and the ontology model is comprehensively reviewed and updated according to a preset period or a specific trigger condition, specifically, the set of all concepts related to web content safety in the ontology model is , is the total number of concepts, for each concept , the semantic feature vector of its current version is , the semantic feature vector of the concept in the new situation obtained after analyzing the new situation is , wherein, is the number of semantic dimensions, and the concept system update necessity index is calculated according to the following formula: , wherein, represents the concept system update necessity index, the greater the value, the greater the gap between the concept system and the requirements in the new situation, and the greater the update necessity, and are the feature values of the concept in the current version and in the new situation in the first semantic dimension, the difference between the two sets of feature values is compared to quantify the change of each concept under the influence of the new situation, and then the update necessity of the entire concept system is comprehensively evaluated.

[0019] Further, in the knowledge extraction step, a joint extraction framework integrating multi-source heterogeneous knowledge is constructed, and a multi-head attention mechanism in deep learning is used to simultaneously focus on key information in different data sources and types. Specifically, for the th data source or type, the feature vector after feature extraction is , wherein represents the feature value of the th data source or type in the th feature dimension, and the query vector , the key vector and the value vector are defined as: , wherein, are the query matrix, the key matrix and the value matrix for the th data source or type, the corresponding vectors are obtained by linear transformation of the original feature vector, and the weight of the th data source or type in knowledge extraction is calculated according to the following formula: , wherein, is the number of heads in the multi-head attention mechanism, is the scaling factor in the th attention calculation, according to the feature vectors of different data sources and types and the calculation of the multi-head attention mechanism, the importance weight of each data source and type in the knowledge extraction process is determined, and the weight obtained by calculation makes the knowledge extraction process pay more attention to those data sources and data types that are more important for web content security analysis.

[0020] Furthermore, in the knowledge fusion step, similarity function and vector space model method are used to calculate the similarity between entities in the extracted knowledge. Specifically, the two entities participating in entity alignment are entity and entities , for entities , define a feature vector set ,in Representing an entity In the feature-level vector description, each vector have Components, that is , for entities , whose feature vector set is ,in , and considering the importance of semantic association between entities in determining whether they are aligned, define a semantic association matrix , whose dimensions are , where the elements Representing an entity In the Feature Levels and Entities In the The semantic correlation between the feature levels ranges from [0, 1], and the entity and entities Entity alignment comprehensive similarity The calculation formula is: ,in, is the weight coefficient.

[0021] Furthermore, in the knowledge fusion step, when entity alignment corresponds to multiple different concepts or attributes, the ambiguity is eliminated by in-depth analysis of semantic relationships and context information using specific algorithms and logic. Specifically, if there is an entity , which has other entities, denoted as With entity Entity alignment comprehensive similarity , for each entity , assuming it has Additional attributes, denoted as , the values ​​of each additional attribute are Define a conflict resolution metric To measure each entity In case of conflict relative to the entity The degree of advantage is calculated as: ,in, and is the weight coefficient and satisfies , used to balance the weight of the part based on the comprehensive similarity exceeding the threshold and the advantage degree based on the additional attribute value in the calculation of the conflict resolution index, For the The weight coefficients of additional attributes, and satisfy , used to weight different additional attributes according to their importance in resolving conflicts, Representative Entity In case of conflict relative to the entity The larger the value, the more advantages the entity has over other competing entities in resolving conflicts after considering the part of the comprehensive similarity exceeding the threshold and the additional attribute value, and the more likely it is to be selected as the final entity. Aligned entities.

[0022] Furthermore, in the monitoring and analysis step, based on the constructed knowledge graph, a specific analysis algorithm and logic are set to conduct in-depth mining and analysis of entities and relationships in the knowledge graph. Specifically, the entity set in the knowledge graph is , the relationship set is , for entities , which has the attribute vector ,in Representing an entity In the The characteristic value on the attribute dimension, For attribute dimensions, for relationships , which has the attribute vector ,in Representing relationships In the The characteristic value on the attribute dimension, For the attribute dimension, define the weight matrix Used to weight entity attributes, its elements Representing an entity In the The weight coefficients on the attribute dimensions define the weight matrix Used to weight relationship attributes, its elements Representing relationships In the The weight coefficient on each dimension is also determined according to the importance of the attribute in the security threat assessment, and the comprehensive assessment value of potential security threats based on the knowledge graph is The calculation formula is: .

[0023] Compared with the existing technology, this web content security processing method based on knowledge graph has the following beneficial effects:

[0024] One, the present application realizes comprehensive monitoring and analysis of web content security by constructing a web content security processing system based on a knowledge graph, and through the use of the reasoning ability of the knowledge graph, including representation learning, correlation analysis, event tracing, behavior prediction, etc., potential security threats, abnormal patterns or correlation in web content can be deeply mined, not only improving the accuracy of security threat identification, but also providing security personnel with an intuitive and easy-to-understand visual interface, facilitating their rapid response and taking appropriate security measures.

[0025] Two, the present application can efficiently collect web content related internet data from multiple sources, including text, image, video and other forms, through the introduction of a multi-modal data fusion adaptive acquisition strategy, not only ensuring the comprehensiveness of data acquisition, but also significantly optimizing the acquisition efficiency and reducing the data processing cost through dynamic adjustment of acquisition frequency and mode, and the use of hierarchical sampling, key frame extraction and other technologies, and the multi-modal data fusion technology makes the collected data have rich correlation information before entering the subsequent processing process, providing a solid foundation for subsequent security analysis.

[0026] Other advantages, objects and features of the present application will be set forth in part in the following specification, and in part will become apparent to those skilled in the art upon examination of the following or will be learned from practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0028] Figure 1 A flow chart of a web content security processing method based on a knowledge graph;

[0029] Figure 2 A knowledge graph diagram of a web content security processing method based on a knowledge graph. DETAILED DESCRIPTION

[0030] The technical solutions in the embodiments of the present application will be described clearly and completely below. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0031] EMBODIMENT

[0032] This embodiment describes in detail a specific application of a knowledge graph-based web content security processing method in processing e-commerce platform web content security threat detection. Through the present invention, the web content security of the e-commerce platform is monitored and managed, various potential security threats and abnormal situations are discovered and handled in a timely manner, and the normal operation of the e-commerce platform and the information security of users are guaranteed.

[0033] First, data related to the web content of the e-commerce platform is collected from multiple key sources. For text information such as product descriptions, user comments, and transaction records on the platform, the text data collection frequency adjustment unit is used to collect them. For example, for the frequently updated description information of popular products (the update frequency is Higher, assuming it is updated 3 times a day), and its importance level coefficient is evaluated The default collection frequency is 0.8 (popular product information has a greater impact on platform operations and user decisions). Once a day, adjustment coefficient Take 0.5, according to the formula , the actual acquisition frequency can be calculated times / day, that is, the system will collect text description information of such popular products at a frequency of about 2.2 times a day. For multimedia content such as product pictures and promotional videos, the collection ratio is determined by the image / video data collection ratio adjustment unit. Assuming that the data volume of a product promotion video is 500MB, resolution 1920x1080 (in terms of horizontal pixel count Vertical pixel count), preset maximum sampling ratio 0.8, reference data size Set to 300MB, reference resolution Set as , adjustment coefficient Take 1, according to the formula , the actual sampling ratio can be calculated , that is, the system will collect data of the promotional video at a ratio of about 23%. The collected text, image, video and other data are fused through the multimodal data fusion unit. Assume that the collected text data vector of a certain product is represented as ( is the feature dimension of text data), and the image data vector after feature extraction is represented as , ( is the feature dimension of image data), and the vector representation of video data after processing is , ( is the feature dimension of video data), set the fusion weight coefficient 、 、 (According to the relative importance of each data type in e-commerce platform content analysis, set) by formula (Where respectively represent the text data vector , image data vector , video data vector The modulus of) fusion, get the fusion data vector , in order to subsequent unified processing.

[0034] In the ontology modeling phase, build the concept system related to e-commerce platform web content security, clearly define a series of concepts, such as "malicious product information" (including false propaganda, infringement, etc.), "sensitive user information disclosure risk" (related to user account, password, payment information, etc. Possible disclosure), "network attack type against e-commerce platform" (such as SQL injection attack, DDoS attack, etc. Attack method against the characteristics of e-commerce platform), "platform security vulnerabilities" (such as payment system vulnerabilities, user authentication vulnerabilities, etc.), etc. Precisely define each concept, for example, "malicious product information" is defined as intentional exaggeration of product efficacy, plagiarism of others' brand image, or dissemination of false authentication information in product description, picture, video, etc. Content, ensure that the concept definition is clear and accurate, avoid ambiguity and ambiguity, at the same time determine the parent-child relationship between the above concepts, when determining the relationship, adopt concept hierarchical relationship determination formula to assist determination, set the two concepts to be determined as concept and concept , The feature similarity index is , at the same time, there is a concept relationship influence factor set , where represents the value of the th influence factor, is the number of influence factors, then the hierarchical relationship tendency value of concept relative to concept is The calculation formula is: , where, is the weight coefficient of the th influence factor, and satisfies , represents the hierarchical relationship tendency value of concept relative to concept , its value range is between [0, 1], when the value of is close to 1, it means that concept is more inclined to be the child class of concept , when the value is close to 0, it means that concept is not the child class of concept A subclass of It's a concept and concept The feature similarity index is calculated as follows: , among which, when When the value of is close, it means that the two concepts are highly similar in features. When the value is close to 0, it means that the similarity of features is extremely low. is the number of feature dimensions, and The concepts and concept In the The characteristic values ​​on the characteristic dimensions are calculated as above, according to the hierarchical relationship tendency value. The size of the data is determined, and the relationship between each subclass and the parent class is clarified. After rigorous logical analysis and sorting, a hierarchical conceptual system is constructed to provide a clear framework for subsequent knowledge organization and processing. For data attribute definition, specific characteristic attributes are given to each e-commerce platform web content security-related concept. For example, for the concept of "malicious product information", attributes such as "false propaganda content type", "infringement type", "occurrence frequency", etc. are given. For each attribute, its specific connotation and definition method are clarified, and key elements such as the attribute value range and data type are specified in detail to ensure the uniqueness and certainty of the attribute definition, so as to facilitate accurate identification and application in subsequent processing. When determining the importance of each characteristic attribute of each concept to the concept in the e-commerce platform web content security analysis, the attribute importance evaluation formula is adopted: Let the concept (such as "malicious product information") feature attributes, denoted as , for each feature attribute ( ), by comparing it with concepts in different network security scenarios The related performance and judgment concept Analyze and evaluate the contribution of relevant safety conditions and other aspects to obtain a comprehensive evaluation value , whose value range is between [0, 1], and there is a set of attribute influencing factors ,in Indicates the The value of the impact factor, is the number of influencing factors, then the characteristic attribute For the concept Importance index The calculation formula is: ,in, For the The weight coefficient of the work, and satisfy , and the uniqueness of each defined data attribute is determined, and the value range, uniqueness, and other constraints are strictly regulated by formulating clear rules for the value range of each attribute and whether it allows repetition, to ensure consistency and compliance of data at the attribute level. After defining the value range of each feature attribute, in order to evaluate whether the set value range is reasonable, i.e., whether it can accurately cover all reasonable values of the attribute in the actual e-commerce platform webpage content security scenario, the attribute value range rationality evaluation formula is used: let the set value range of feature attribute ( such as "false propaganda content type" ) be , and through statistical analysis of a large number of historical network security data, related standards and specifications, and actual application cases of the value of this attribute, an actual value distribution function is obtained, where represents the value of attribute , and the attribute value range rationality index is calculated by the formula . When determining whether each feature attribute satisfies the uniqueness constraint condition, the attribute uniqueness constraint evaluation formula is used: let the concept ( such as "malicious product information" ) have feature attribute ( such as "false propaganda content type" ), and in a sample set containing e-commerce platform webpage content security related instances ( such as different product information records, security detection events, etc. ), the value of attribute in these instances is counted, let the value of attribute in the instance be ( , then the uniqueness constraint index of attribute is calculated by the formula , where is a Kronecker function, when , , , , , represents the number of combinations of 2 elements selected from elements, and the calculation formula is ​,For the definition of semantic relationships, define a variety of semantic relationships related to the security of e-commerce platform web page content. For each semantic relationship, clarify its specific meaning and judgment criteria, and elaborate on the conditions under which a certain semantic relationship can be considered to be established, to ensure the clarity and operability of the semantic relationship definition, and provide an accurate basis for subsequent knowledge reasoning and analysis. For example, for the association between "malicious product information" and "sensitive user information leakage risk", the judgment criteria can be set as: when there is false propaganda in the product information and it induces users to input additional information, and the transaction process to which the product belongs involves the transmission of key user information, this association is considered to be established. For the formulation of naming standards, formulate unified and clear naming standards, which are applicable to concepts, attributes and relationships related to the security of e-commerce platform web page content, and stipulate the naming format, rules and principles to be followed to ensure that each element The naming is consistent and clear, which is convenient for subsequent understanding, operation and maintenance in the process of knowledge graph construction, query and management. For the establishment of axioms and rules, a series of axioms and rules are established. These axioms and rules are determined based on the business logic and actual needs of the e-commerce platform web content security. The specific triggering conditions and execution logic of each axiom and rule are clarified, and the corresponding actions should be triggered under what conditions are specified in detail to ensure that in the entire e-commerce platform web content security processing process, the operations and decisions of each link have rules to follow and evidence to rely on. At the same time, a dynamic update mechanism is established to comprehensively review and update the entire ontology model according to the preset update cycle or specific trigger conditions. When conducting a comprehensive review, the concept system update necessity assessment formula is used to determine whether the concept system needs to be updated. Suppose the set of all concepts related to the e-commerce platform web content security in the ontology model is ,in is the total number of concepts, for each concept , let the semantic feature vector of the current version be After analyzing the new situation, the semantic feature vector of the concept in the new context is ,in, is the number of semantic dimensions, then the conceptual system update necessity index The calculation formula is: ,in, Represents the necessity index of updating the concept system. The larger its value, the greater the gap between the concept system and the requirements in the new situation, and the greater the necessity of updating. and The concepts In the current version and new situation By comparing the differences between the two sets of feature values, we can quantify the degree of change of each concept under the influence of the new situation, and then comprehensively evaluate the necessity of updating the entire concept system.

[0035] In the knowledge extraction stage, for the rich commodity description text on the e-commerce platform, natural language processing technology is used for knowledge extraction. First, the text is preprocessed, including word segmentation, part-of-speech tagging, and removing stop words. Based on the preprocessed text, specific rules and patterns are set to identify key information. For commodity description, focus on information related to product characteristics, functions, advantages, and expressions that may involve safety hazards. Use named entity recognition (NER) technology to identify various entities in the text, such as product names, brands, and models as specific entities, and "malicious commodity information" and "sensitive user information" as abstract safety-related entities. For image data such as commodity pictures and promotional posters on the e-commerce platform, computer vision technology is used for feature extraction. For example, use a convolutional neural network (CNN) model to process images and extract color features, texture features, and shape features. Based on the extracted image features, further identify safety-related features, such as unclear product identification, suspected tampered authentication marks, and exaggerated visual effects that may imply false advertising (such as excessively enlarged product effect displays). These features are extracted as safety-related information. These feature information will be associated with the corresponding commodity entity in the subsequent steps to reflect the relationship between image data and safety conditions in the knowledge graph. The image features corresponding to the same commodity entity identified in the text data are associated. According to the content presented in the image and the echo of the text description, the relationship between entities is constructed. For video data such as commodity promotional videos on the e-commerce platform, first divide the video into multiple video frames, then analyze each video frame. By analyzing the video frames, features such as the appearance of the phone from different angles, the user's hand gestures when operating the phone, and the user's satisfied expression on the face can be extracted. Based on the analysis of the video frames, safety-related information is captured. The information extracted from the video data is integrated with the information extracted from the text and image data. Through the commodity identification, voiceover, and other information in the video, the corresponding commodity entity is found. The relationships presented in the video are merged with the relationships constructed in the text and image data to form a complete knowledge system about the commodity and related safety conditions, so that the knowledge graph can accurately present the information in the subsequent steps.

[0036] When performing knowledge fusion, to determine whether two entities should be aligned, an improved entity alignment comprehensive similarity formula is used to calculate the comprehensive similarity between entities. Let the two entities participating in entity alignment be entity and entity For entity , describe its features from multiple aspects and define a feature vector set , where represent entities In the feature layer of the first characteristic vector, each vector has components, that is , for an entity , the feature vector set is , wherein , the feature layer includes the basic attributes of the entity (such as name, type, etc.), associated information (such as the association relationship with other entities, the category to which it belongs, etc.), behavioral characteristics (such as operation behavior in a specific scenario, activity regularity, etc.), and the like. Through detailed analysis and quantification of the entity, the vector description of each layer is obtained. For an entity , in the "basic attribute" feature layer, contains the character feature vector of the product name and the classification code of the product type and the like. In the "associated information" feature layer, contains the user evaluation vector related to the product and the code of the region of the e-commerce platform to which it belongs. In the "behavioral characteristics" feature layer, contains the sales trend vector of the product and the promotion activity frequency vector and the like. Similarly, for an entity , the feature vector set is , wherein , in addition, considering the importance of semantic association between entities in determining whether to align, a semantic association matrix is defined, which has a dimension of , wherein the element represents the semantic association degree between the entity in the first feature layer and the entity in the first feature layer, and the value range is between [0, 1]. Then, the entity alignment comprehensive similarity between the entity and the entity is calculated according to the following formula: , wherein is a weight coefficient, wherein is a weight coefficient, and satisfies , which is used to weight according to the importance of the association between different feature layers and the proportion of the semantic association degree in the comprehensive similarity calculation. If the calculated is greater than a preset alignment threshold , it is considered that the entity and the entity can be aligned, and they are merged into one entity, so as to unify the processing of related knowledge information in the knowledge fusion process, avoid knowledge redundancy, and improve the accuracy of subsequent analysis and processing.

[0037] When there are multiple entities that reach or approach the alignment threshold in the integrated similarity with a certain entity in the entity alignment process, but there are differences between these entities, the final alignment selection is determined by the entity alignment conflict resolution formula to resolve the conflict. Let there be an entity , which has other entities, denoted as The entity alignment integrated similarity of entity with entity , for each entity , let it have additional attributes, denoted as , and the value of each additional attribute is Define a conflict resolution indicator to measure the advantage of each entity over entity in the conflict situation, the calculation formula is: , where and are weight coefficients, and satisfy , used to balance the proportion of the part based on the integrated similarity exceeding the threshold and the advantage degree based on the additional attribute value in the conflict resolution indicator calculation, is the weight coefficient of the additional attribute, and satisfies , used to weight according to the importance of different additional attributes in resolving conflicts, represents the conflict resolution indicator of entity over entity in the conflict situation, the greater the value, the more advantageous the entity is in resolving conflicts relative to other competing entities, and the easier it is to be selected as the final entity aligned with entity .

[0038] After the completion of knowledge fusion, the construction of knowledge graph is carried out, taking various entities determined in the ontology modeling and knowledge fusion stage as the nodes of the knowledge graph, for example, "malicious commodity information", "sensitive user information leakage risk", "network attack type of e-commerce platform", "platform security vulnerability" and the like. Each node corresponds to a node in the graph. For each node, in addition to containing the basic concept information, the relevant attribute information determined in the previous stage is also associated with the node, such as the "malicious commodity information" node, in addition to labeling its name, it will also associate "false propaganda content type", "infringement type", "occurrence frequency" and the like. According to the semantic relationship defined in the ontology modeling stage, edges are created between the corresponding nodes to represent these relationships, for example, if there is an association relationship between "malicious commodity information" and "sensitive user information leakage risk", a directed edge is created between the nodes representing the two entities. The direction of the edge can be determined according to the logical direction of the relationship. Similarly, for the corresponding relationship between "network attack type of e-commerce platform" and "platform security vulnerability", a suitable edge is also created between their corresponding nodes. In this way, various entity nodes are connected through semantic relationship edges to form a complete knowledge graph structure. The constructed knowledge graph is stored in a graph database. The graph database can efficiently handle the storage and query operations of nodes and edges, and is suitable for the storage requirements of knowledge graphs with complex relationship structures. In the storage process, the metadata of the graph is standardized to facilitate subsequent query, update and maintenance operations, for example, a unique identifier is set for each node and edge to facilitate quick positioning and operation of the corresponding elements in the database.

[0039] After the knowledge graph is constructed, monitoring analysis is performed to monitor and analyze the content of the e-commerce platform webpage in real time or periodically to discover potential security threats and abnormal situations. The latest webpage content related data is continuously collected from various data sources of the e-commerce platform, including newly published commodity information, user comments, transaction records, platform system logs, etc. These newly collected data are processed according to the previous data processing procedure, and then updated to the attributes of the corresponding nodes and edges in the knowledge graph. For example, if the user comments of a certain commodity newly collected mention suspected false propaganda content, the "occurrence frequency" attribute of the "malicious commodity information" node is correspondingly increased, and the value of the "false propaganda content type" attribute may be further updated according to the comment content. The updated knowledge graph is detected for abnormalities using pre-set rules and models. The rules established in the axiom and rule setting stage are used for detection, and deep learning models are used to analyze the data in the knowledge graph. For example, a classification model is trained to determine whether the newly collected commodity information belongs to the category of "malicious commodity information". The various features of the commodity information (such as text description, picture features, sales data, etc.) are used as inputs of the model, and the model makes a judgment according to the previously trained parameters and algorithms. If the model output indicates that the commodity information has a high possibility of belonging to "malicious commodity information", the corresponding node in the knowledge graph is also marked as an abnormal situation, and the nodes and edges in the knowledge graph are analyzed to discover complex relationships and potential risks hidden in the data. Specifically, let the entity set in the knowledge graph be , the relationship set be , the entity have attribute vector , where represents the feature value of the entity in the th attribute dimension, and is the attribute dimension. The relationship has attribute vector , where represents the feature value of the relationship in the th attribute dimension, and is the attribute dimension. The weight matrix is defined to weight the entity attributes, and its element represents the weight coefficient of the entity in the th attribute dimension. The weight matrix is defined to weight the relationship attributes, and its element represents the weight coefficient of the relationship in the The weight coefficient on the action dimension is also determined according to the importance of the attribute in the security threat evaluation, and then a comprehensive evaluation value of the potential security threat based on the knowledge graph is calculated The calculation formula is as follows: If it is found that the user's purchase behavior has a significant downward trend after the emergence of a certain type of malicious commodity information, and there is a time correlation with the ongoing promotion activities of the platform, it can be inferred that the malicious commodity information may have a negative impact on the platform promotion activities, and further analysis of the reasons and corresponding measures are taken,

[0040] Once abnormal conditions or valuable analysis results are found in the monitoring and analysis stage, the result display and response stage is entered to timely show information to relevant personnel and take effective measures. By developing a visual user interface, the knowledge graph and the results of monitoring and analysis are displayed in an intuitive graphical manner. On the interface, each node (representing different entities), edge (representing semantic relationship) and node or area marked as abnormal condition can be clearly seen. For example, the "malicious commodity information" node is highlighted in red, and if the associated "sensitive user information leakage risk" node is also in an abnormal state, the node is marked with a flashing red border so that users can see the problem area at a glance. At the same time, the change of the relevant attribute value is displayed through the chart (such as bar chart, line chart, etc.), such as the curve of the frequency of "malicious commodity information" over time, so that users can more intuitively understand the development trend of the problem.

[0041] In summary, through the above steps, the e-commerce platform web content security can be comprehensively and systematically monitored and managed, potential security threats and abnormal conditions can be timely discovered and handled, and the normal operation of the e-commerce platform and the information security of the users can be ensured.

[0042] It will be apparent to those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, but can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be considered exemplary and non-limiting, and the scope of the present application is defined by the appended claims, not the above description, and therefore all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be considered as limiting the claims involved.

Claims

1. A webpage content security processing method based on knowledge graph, characterized in that: The method comprises the following specific steps: Data collection: Collect web-related Internet content data in the form of text, images, and videos from network service providers, internal networks of enterprises or organizations, and website servers. Develop adaptive collection frequencies and methods based on the characteristics of data from different sources and the needs of web content security analysis. For text data, set its update frequency to , the importance level coefficient is , the preset basic acquisition frequency is , the adjustment coefficient is , its actual acquisition frequency The calculation formula is: , for image or video data, let the data size be , with a resolution of , the preset maximum sampling ratio is , the reference data size is , the reference resolution is , the adjustment coefficient is , its actual sampling ratio The calculation formula is: , and use multimodal data fusion technology to perform real-time fusion processing on different modal data during collection, so that the collected data can meet the requirements of subsequent analysis; Ontology modeling: Clearly define concepts related to web content security and precisely define their scope and characteristics. Identify parent-child relationships between concepts to build a hierarchical conceptual system. Each concept is assigned specific characteristic attributes, clarifying its connotation, definition method, value range, and uniqueness constraints. Furthermore, multiple semantic relationships are defined, with clear meanings and judgment criteria. A unified naming convention is developed, along with axioms and rules and clear triggering conditions and execution logic. Furthermore, a semantic expansion and dynamic update mechanism is established to automatically expand the semantics of related concepts based on new situations and review and update the ontology model according to preset cycles or specific triggering conditions. Knowledge extraction: Using natural language processing technology, we pre-process different types of data to ensure consistency in format and feature representation. By building a joint extraction framework that integrates multi-source heterogeneous knowledge and leveraging the multi-head attention mechanism in deep learning to simultaneously focus on key information from different data sources and types, we can accurately extract entity, relationship, and attribute information, providing a high-quality knowledge foundation for subsequent knowledge fusion and graph construction. Knowledge fusion: Similarity functions and vector space model methods are used to calculate the similarity between entities in the extracted knowledge. Entities that meet the set similarity threshold are aligned and merged and treated as the same entity. When entity alignments are ambiguous, corresponding to multiple different concepts or attributes, specific algorithms and logic are used to eliminate ambiguity through in-depth analysis of semantic relationships and contextual information. Graph construction: Select a graph database and accurately import the entities and inter-entity relationships after knowledge fusion processing into the graph database according to their format requirements through a specific input interface and data conversion process to build a knowledge graph, forming a semantic network containing entities and inter-entity relationships, and clarifying the storage and representation methods of each entity and relationship in the knowledge graph; Monitoring and analysis: Based on the constructed knowledge graph, specific analysis algorithms and logic are set to conduct in-depth mining and analysis of entity relationships in the knowledge graph to explore potential security threats, abnormal patterns or associations; Result display and response: Use visualization technology and tools to design layout, display methods and interactive functions to display knowledge graph relationships and monitoring and analysis results, so that security personnel can take appropriate measures based on the results.

2. A webpage content security processing method based on knowledge graph according to claim 1, characterized in that: In the data collection step, multimodal data fusion technology is used to fuse different modal data in real time during collection. Specifically, the collected text data vector is represented as , is the feature dimension of text data, and the vector representation of image data after feature extraction is , is the feature dimension of image data, and the vector representation of video data after processing is , is the feature dimension of video data, which is fused and the fusion weight coefficient is set 、 ,γ and satisfy , the vector after multimodal data fusion The calculation formula is: ,in, 、 、 Represents text data vectors respectively , image data vector , video data vector Model.

3. A webpage content security processing method based on knowledge graph according to claim 1, characterized in that: In the ontology modeling step, the parent-child relationship between concepts is determined to build a hierarchical concept system. The two concepts whose relationship is to be determined are defined as concepts. and concepts , The feature similarity index is , and suppose there is a concept relationship impact factor set ,in Indicates the The value of the impact factor, is the number of impact factors, then the concept Relative to the concept Hierarchical relationship tendency value The calculation formula is: ,in, For the The weight coefficient of the impact factor, and satisfy , Representative Concept Relative to the concept The hierarchical relationship tendency value of When the value is close to 1, it means that the concept More like a concept When the value is close to 0, it means the concept Not a concept A subclass of It's a concept and concept feature similarity index.

Citation Information

Patent Citations

  • Knowledge-driven business operation graph construction method

    CN112507136A

  • Analysis processing and knowledge graph construction method based on multi-source heterogeneous big data

    CN115640406A