Method, apparatus and electronic device for identifying associated abnormal data set based on big data
By constructing enterprise structured and unstructured data sets and using graph models to identify abnormal areas of association density, the problem of insufficient accuracy in identifying abnormal data of complex enterprise relationships in the existing technology is solved, and efficient and automatic abnormal data recognition and analysis is achieved.
Patent Information
- Application Number
- CN202510185905.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-02-20
Smart Images

Figure CN119669989B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and particularly to a method, apparatus, and electronic device for identifying associated abnormal data sets based on big data. Background Art
[0002] With the rapid development of information technology, enterprises and organizations have accumulated a large amount of data resources. These data contain rich information, but at the same time, they also bring problems of data management and security. How to effectively manage and use this data to ensure the security and accuracy of the data has become an important research topic. Existing data anomaly detection technologies have made certain progress, but still face many challenges in the face of large-scale, multi-source heterogeneous data.
[0003] Currently, existing data anomaly detection technologies often rely on a single data type or simple statistical models, and it is difficult to meet the actual needs of diverse data types and complex association relationships. This limitation leads to the inability to accurately identify real abnormal data in practical applications, especially when facing complex upstream and downstream relationships between enterprises, thus affecting the accuracy of abnormal data detection.
[0004] Therefore, there is an urgent need for a method, apparatus, and electronic device for identifying associated abnormal data sets based on big data. Summary of the Invention
[0005] This application provides a method, apparatus, and electronic device for identifying associated abnormal data sets based on big data, which improves the accuracy of abnormal data monitoring.
[0006] In the first aspect of this application, a method for identifying an associated abnormal data set based on big data is provided. The method includes: obtaining a target enterprise and its associated upstream and downstream enterprises; constructing an original data set according to the target enterprise, the upstream enterprise, and the downstream enterprise, where the original data set includes structured data and unstructured data; preprocessing the structured data, performing triple conversion on the unstructured data, and merging the preprocessed structured data and the converted unstructured data to obtain a target data set; constructing a data association graph based on the target data set, where the nodes in the data association graph represent data elements in the target data set, the edges in the data association graph represent the association relationships between the data elements, and the association relationships are obtained by calculating the similarity between the data elements; in the data association graph, identifying an associated density abnormal area, where the associated density abnormal area is an area with abnormal associated density, and the associated density refers to the number of edges in a unit area; marking the data elements included in the associated density abnormal area as abnormal data; and determining the corresponding abnormal data set according to the position information of the abnormal data in the original data set.
[0007] By adopting the above technical solution, data of a target enterprise, its upstream enterprises and downstream enterprises are obtained to construct an original data set containing structured data and unstructured data, and the original data set is preprocessed and transformed to obtain a target data set. A data association graph is constructed based on the target data set, where nodes represent data elements and edges represent the association relationships between data elements. Association density abnormal regions are identified in the data association graph, data elements within the association density abnormal regions are marked as abnormal data, and corresponding abnormal data sets are determined according to the position information of the abnormal data in the original data set. This method makes full use of the structured data and unstructured data of the target enterprise and its upstream and downstream enterprises. Through data preprocessing and transformation, heterogeneous data are uniformly represented, and a comprehensive and high-quality target data set is constructed. A graph model is used to represent the association relationships between data elements, and the association strength is quantified by similarity, so as to depict the complex associations between data elements. Based on the graph model, association density abnormal regions are identified, and abnormal data sets can be automatically and efficiently discovered. Mapping the abnormal data back to the original data set can determine the source and context of the abnormal data set, facilitating subsequent analysis and disposal. Therefore, this method can make full use of big data to automatically, accurately and efficiently identify associated abnormal data sets, providing data support for applications such as enterprise risk prevention and public opinion monitoring.
[0008] Optionally, identifying the association density abnormal regions in the data association graph specifically includes: performing grid processing on the data association graph to divide the data association graph into multiple grid units of the same size, where each grid unit includes at least one node; calculating the association density of each grid unit, where the association density is the actual number of edges within the grid unit divided by a preset number, and the preset number is the maximum number of edges that may exist within the grid unit; according to the association density, calculating the domain association density of a first grid unit, where the domain association density is the average of the association densities of second grid units adjacent to the first grid unit, and the first grid unit is any one of the multiple grid units; calculating the absolute value of the difference between the association density of the first grid unit and the domain association density of the first grid unit, and taking the absolute value as the density difference degree of the first grid unit; determining the first grid unit with the density difference degree greater than or equal to a preset difference degree threshold as an abnormal grid unit, and merging adjacent abnormal grid units into the association density abnormal region.
[0009] By adopting the above technical solution, through grid processing of the data association graph, the data association graph is divided into grid cells of the same size, and each grid cell contains at least one node. Calculate the association density of each grid cell, that is, the actual number of edges divided by the maximum possible number of edges. Calculate the domain association density of each grid cell, that is, the average value of the association densities of adjacent grid cells. Calculate the absolute value of the difference between the association density and the domain association density of each grid cell as the density difference degree of the grid cell. Mark the grid cells with a density difference degree greater than or equal to the preset difference degree threshold as abnormal cells, and merge adjacent abnormal cells into an association density abnormal region. This method cleverly uses grid processing to transform the irregular graph structure into a regular grid structure, which is convenient for calculating and comparing the association density. By calculating the association density and the domain association density of the grid cells, the degree of association tightness inside and around the cells can be measured. Using the density difference degree as an index, the abnormality degree of the grid cells can be quantified, and the abnormal grid cells can be automatically identified. By adopting the method of region merging, a complete and continuous abnormal region can be obtained. Therefore, this method can effectively identify the association density abnormal region and provide a reliable basis for the identification of abnormal data sets.
[0010] Optionally, constructing the original data set according to the target enterprise, the upstream enterprise, and the downstream enterprise specifically includes: obtaining the structured data of the target enterprise, the upstream enterprise, and the downstream enterprise, where the structured data includes enterprise basic information data, financial data, and transaction data; obtaining the unstructured data of the target enterprise, the upstream enterprise, and the downstream enterprise, where the unstructured data includes enterprise news data, judicial litigation data, and patent data; performing data cleaning and data integration on the structured data and the unstructured data to obtain the original data set.
[0011] By adopting the above technical solution, by obtaining the structured data and unstructured data of the target enterprise and its upstream and downstream enterprises, and performing data cleaning and data integration on these two types of data respectively, the original data set is obtained. The structured data includes enterprise basic information, financial data, transaction data, etc., and the unstructured data includes enterprise news, judicial litigation, patent data, etc. This method comprehensively collects multi-source heterogeneous data inside and outside the enterprise, covering multiple business fields and data types of the enterprise. Through data cleaning, low-quality data such as noise and errors are removed, improving the accuracy and reliability of the data. Through data integration, data from different sources and formats are unified and associated to form a complete and consistent original data set. This method lays a solid data foundation for the identification of associated abnormal data sets and helps to improve the accuracy and comprehensiveness of the identification.
[0012] Optionally, the step of converting the unstructured data into triples specifically includes: performing natural language processing on the unstructured data to extract key information; converting the key information into triples according to a preset template, where the triples include a subject, a predicate, and an object; and using the triples as data elements of the target data set.
[0013] By adopting the above technical solution, through performing natural language processing on unstructured data, key information is extracted, and the key information is converted into triples according to a preset template. The triples include three elements: a subject, a predicate, and an object. Using the triples as data elements of the target data set, they participate in subsequent analysis together with the structured data. This method cleverly utilizes natural language processing technology to extract structured semantic information from unstructured data, revealing the internal logical relationships of the text. Adopting the triple representation form can cover most types of semantic relationships and has strong expressive power. Converting unstructured data into triples makes it have the same representation form as structured data, facilitating the fusion and correlation analysis of the two types of data. Based on the unstructured data represented by triples, it can enrich the semantic information of the target data set, improving the coverage and accuracy of the correlation relationships. Therefore, this method provides strong support for the fusion and analysis of unstructured data, expanding the applicable scope of identifying associated abnormal data sets.
[0014] Optionally, constructing a data association graph based on the target data set specifically includes: using each data element in the target data set as a node of the data association graph; calculating the similarity between a first data element and a second data element, where the first data element and the second data element are any two of the multiple data elements; and if it is determined that the similarity is greater than or equal to a preset similarity threshold, establishing an edge between the first data element and the second data element to obtain the data association graph.
[0015] By adopting the above technical solution, using the data elements in the target data set as nodes of the graph, calculating the similarity between any two data elements, and if the similarity is greater than or equal to the preset similarity threshold, establishing an edge between the two data elements. This method intuitively and vividly depicts the association relationships between data elements, and objectively and accurately measures the strength of the association through the quantitative index of similarity. The setting of the preset similarity threshold can filter out weakly associated or unassociated data element pairs, making the association graph more concise and reliable. Using the data association graph as the basis for abnormal region identification can make full use of the association information between data elements to depict the association pattern and abnormal pattern of the data set from both the global and local perspectives. The quality and accuracy of the data association graph directly affect the effect of abnormal region identification. This method provides key data structures and algorithm support for abnormal region identification, helping to improve the effect and efficiency of identification.
[0016] Optionally, calculating the similarity between the first data element and the second data element specifically includes: extracting a first feature vector of the first data element and extracting a second feature vector of the second data element; calculating a cosine similarity between the first feature vector and the second feature vector, and using the cosine similarity as the similarity between the first data element and the second data element.
[0017] By adopting the above technical solution, by extracting the feature vectors of any two data elements and calculating the cosine similarity of the two feature vectors as the similarity of the two data elements. This method cleverly utilizes the vector space model to map the data elements into a high-dimensional feature space and characterizes the features of the data elements through the feature vectors. Using the cosine similarity as the measurement standard for the similarity of data elements can overcome the influence brought by differences such as data types and numerical ranges, and provide a unified and reliable similarity evaluation. The similarity calculation method based on feature vectors and cosine similarity provides a reasonable and effective similarity measurement for the construction of the data association graph, which helps to improve the accuracy and interpretability of the association relationship.
[0018] Optionally, determining the corresponding abnormal data set according to the position information of the abnormal data in the original data set specifically includes: obtaining the position information of the abnormal data in the target data set, where the position information includes the row number and column number of the abnormal data in the target data set; obtaining the position information of the abnormal data in the structured data and the unstructured data according to the position information of the abnormal data in the target data set; determining the target structured data associated with the abnormal data according to the position information of the abnormal data in the structured data, and determining the target unstructured data associated with the abnormal data according to the position information of the abnormal data in the unstructured data; merging the target structured data and the target unstructured data to obtain an original data subset; using the original data subset as the abnormal data set corresponding to the abnormal data.
[0019] By adopting the above technical solution, the position information of abnormal data in the target dataset is obtained, including the row number and column number. Then, according to the position information, the positions of the abnormal data in the structured data and the unstructured data are obtained. According to the positions of the abnormal data in the structured data and the unstructured data, the target structured data and the target unstructured data associated with the abnormal data are determined, and they are merged to obtain the original data subset, which is used as the abnormal dataset corresponding to the abnormal data. This method makes full use of the position correspondence relationship of abnormal data in different datasets. Through position mapping, the abnormal data is restored from abstract data elements to specific structured records and unstructured texts. By extracting the associated data of the abnormal data, an abnormal dataset is formed. Merging the structured data and the unstructured data can analyze abnormal situations from multiple dimensions and levels, and mine the associated features and evolution laws of the abnormalities. The method of locating and extracting the abnormal dataset based on the original dataset provides rich and reliable data support for the interpretation and disposal of abnormalities, and helps to analyze the causes of abnormalities, evaluate the impacts of abnormalities, and formulate countermeasures.
[0020] In the second aspect of the present application, a device for identifying an associated abnormal dataset based on big data is provided, which is characterized in that the device includes an acquisition module and a processing module, wherein: the acquisition module is used to acquire a target enterprise and upstream and downstream enterprises associated with the target enterprise; the processing module is used to construct an original dataset according to the target enterprise, the upstream enterprise, and the downstream enterprise, and the original dataset includes structured data and unstructured data; the processing module is further used to preprocess the structured data, perform triple transformation on the unstructured data, and merge the preprocessed structured data and the transformed unstructured data to obtain a target dataset; the processing module is further used to construct a data association graph based on the target dataset, where the nodes in the data association graph represent data elements in the target dataset, and the edges in the data association graph represent the association relationships between the data elements, and the association relationships are obtained through similarity calculation between the data elements; the processing module is further used to identify an associated density abnormal area in the data association graph, and the associated density abnormal area is an area where the associated density is abnormal, and the associated density refers to the number of edges in a unit area; the processing module is further used to mark the data elements included in the associated density abnormal area as abnormal data; the processing module is further used to determine the corresponding abnormal dataset according to the position information of the abnormal data in the original dataset.
[0021] In a third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. Both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory, so that the electronic device executes the method described in any one of the above.
[0022] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions, and when the instructions are executed, the method described in any one of the above is executed.
[0023] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0024] 1. By obtaining data of a target enterprise and its upstream and downstream enterprises, constructing an original data set including structured data and unstructured data, and performing preprocessing and transformation on the original data set, a target data set is obtained. Based on the target data set, a data association graph is constructed, where nodes represent data elements and edges represent the association relationships between data elements. In the data association graph, regions with abnormal association density are identified, the data elements within the regions with abnormal association density are marked as abnormal data, and according to the position information of the abnormal data in the original data set, the corresponding abnormal data set is determined. This method makes full use of the structured data and unstructured data of the target enterprise and its upstream and downstream enterprises. Through data preprocessing and transformation, heterogeneous data is uniformly represented, and a comprehensive and high-quality target data set is constructed. The graph model is used to represent the association relationships between data elements, and the association strength is quantified by similarity, thereby depicting the complex associations between data elements. Based on the graph model, regions with abnormal association density are identified, which can automatically and efficiently discover abnormal data sets. Mapping the abnormal data back to the original data set can determine the source and context of the abnormal data set, facilitating subsequent analysis and disposal. Therefore, this method can make full use of big data to automatically, accurately, and efficiently identify associated abnormal data sets, providing data support for applications such as enterprise risk prevention and public opinion monitoring.
[0025] 2. By performing grid processing on the data association graph, the data association graph is divided into grid cells of the same size, and each grid cell contains at least one node. Calculate the association density of each grid cell, that is, the actual number of edges divided by the maximum possible number of edges. Calculate the domain association density of each grid cell, that is, the average value of the association densities of adjacent grid cells. Calculate the absolute value of the difference between the association density and the domain association density of each grid cell as the density difference degree of the grid cell. Mark the grid cells with a density difference degree greater than or equal to the preset difference degree threshold as abnormal cells, and merge adjacent abnormal cells into an association density abnormal region. This method cleverly uses grid processing to transform the irregular graph structure into a regular grid structure, which is convenient for calculating and comparing the association density. By calculating the association density and the domain association density of the grid cells, the degree of association tightness inside and around the cells can be measured. Using the density difference degree as an index, the abnormality degree of the grid cells can be quantified, and abnormal grid cells can be automatically identified. By adopting the method of region merging, a complete and continuous abnormal region can be obtained. Therefore, this method can effectively identify the association density abnormal region and provide a reliable basis for the identification of abnormal data sets.
[0026] 3. By obtaining the structured data and unstructured data of the target enterprise and its upstream and downstream enterprises, perform data cleaning and data integration on these two types of data respectively to obtain the original data set. Structured data includes enterprise basic information, financial data, transaction data, etc., and unstructured data includes enterprise news, judicial litigation, patent data, etc. This method comprehensively collects multi-source heterogeneous data inside and outside the enterprise, covering multiple business fields and data types of the enterprise. Through data cleaning, low-quality data such as noise and errors are removed, improving the accuracy and reliability of the data. Through data integration, data from different sources and formats are unified and associated to form a complete and consistent original data set. This method lays a solid data foundation for the identification of associated abnormal data sets and helps to improve the accuracy and comprehensiveness of the identification. Description of the Drawings
[0027] Figure 1 is a schematic flowchart of a method for identifying an associated abnormal data set based on big data disclosed in an embodiment of the present application;
[0028] Figure 2 is a schematic module diagram of a device for identifying an associated abnormal data set based on big data disclosed in an embodiment of the present application;
[0029] Figure 3 is a schematic structural diagram of an electronic device disclosed in an embodiment of the present application.
[0030] Description of the drawing reference numerals: 201, acquisition module; 202, processing module; 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed implementation mode
[0031] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments.
[0032] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "for example" or "for instance" is intended to present relevant concepts in a specific way.
[0033] In the description of the embodiments of this application, the meaning of the term "plurality" refers to two or more. For example, a plurality of systems refers to two or more systems, and a plurality of screen terminals refers to two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the technical features indicated. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0034] This application provides a method for identifying associated abnormal data sets based on big data. Refer to Figure 1 , Figure 1 is a schematic flowchart of the method for identifying associated abnormal data sets based on big data provided in the embodiments of this application. This method is applied to a server. The server is a server that executes a program for identifying associated abnormal data sets based on big data. The server can be a single server, or a server cluster composed of multiple servers, or a cloud computing service center. This method includes steps S101 to S107, and the above steps are as follows:
[0035] Step S101: Obtain the target enterprise and the upstream and downstream enterprises associated with the target enterprise.
[0036] In step S101, the server can connect to the databases of various enterprises and directly extract the information of the target enterprise, upstream enterprises, and downstream enterprises from the databases. The databases of enterprises usually contain the basic information of the enterprises, supplier information, customer information, etc. The server can determine the target enterprise and its associated enterprises (upstream enterprises and downstream enterprises) based on this information.
[0037] Step S102: Construct an original data set based on the target enterprise, upstream enterprises, and downstream enterprises. The original data set includes structured data and unstructured data.
[0038] In step S102, an original data set is constructed based on the target enterprise, upstream enterprises, and downstream enterprises, specifically including: obtaining the structured data of the target enterprise, upstream enterprises, and downstream enterprises. The structured data includes enterprise basic information data, financial data, and transaction data; obtaining the unstructured data of the target enterprise, upstream enterprises, and downstream enterprises. The unstructured data includes enterprise news data, judicial litigation data, and patent data; performing data cleaning and data integration on the structured data and unstructured data to obtain the original data set.
[0039] Specifically, the server constructs an original data set based on the target enterprise, upstream enterprises, and downstream enterprises. The original data set includes structured data and unstructured data.
[0040] For the structured data, the server obtains the enterprise basic information data, financial data, and transaction data of the target enterprise, upstream enterprises, and downstream enterprises. The enterprise basic information data includes enterprise name, registered address, registered capital, founding time, business scope, etc.; the financial data includes the balance sheet, income statement, cash flow statement, etc. of the enterprise; the transaction data includes the purchase orders, sales orders, logistics documents, etc. of the enterprise. The server obtains this structured data through multiple channels such as connecting to the internal databases of enterprises, third-party data service providers, and public databases of government departments.
[0041] For the unstructured data, the server obtains the enterprise news data, judicial litigation data, and patent data of the target enterprise, upstream enterprises, and downstream enterprises. The enterprise news data includes news reports, announcement information, etc. related to the enterprise; the judicial litigation data includes the legal litigation records in which the enterprise participates as the plaintiff, defendant, or third party; the patent data includes the patent information applied for or owned by the enterprise. The server can obtain this unstructured data by crawling news websites, judicial case databases, patent databases, etc. on the Internet.
[0042] After obtaining structured data and unstructured data, the server cleans and integrates the data. Data cleaning refers to identifying and correcting errors, inconsistencies, and duplication issues in the data to improve data quality. Data integration refers to integrating data from multiple sources to form a unified and coherent data view. For example, the server obtains a company's financial data from different databases, which may have different data formats and field names. The server maps this data into a unified data model. Another example is that the server crawls a large amount of company news from news websites and needs to associate it with existing structured data. For example, it matches the company names mentioned in the news with the company names in the company basic information data and adds the news information to the corresponding company records. Through data cleaning and integration, the server finally obtains an original dataset that combines structured data and unstructured data.
[0043] Step S103: Preprocess the structured data, perform triple transformation on the unstructured data, and merge the preprocessed structured data and the transformed unstructured data to obtain a target dataset.
[0044] In step S103, the steps of performing triple transformation on the unstructured data specifically include: performing natural language processing on the unstructured data to extract key information; according to a preset template, converting the key information into triples, where the triples include a subject, a predicate, and an object; and using the triples as data elements of the target dataset.
[0045] Specifically, for structured data, the server performs preprocessing operations such as data cleaning and format conversion. Specifically: The server first performs a data quality check on the structured data to identify and handle data quality issues such as missing values, outliers, and duplicate values. For example, for missing values, the server can use methods such as mean filling and median filling to fill them according to the field type and distribution characteristics; for outliers, the server uses methods such as truncation, replacement, and deletion according to the field value range and distribution characteristics; for duplicate values, the server uses methods such as deduplication, merging, and retention according to the uniqueness constraint of the field. Then, the server performs format conversion on the structured data, converting data from different sources and different formats into a unified format. For example, the server converts CSV format data into JSON format, converts date strings into standard date and time formats, and converts string-type numerical values into numerical types, etc.
[0046] The server performs triple conversion on unstructured data, converting the unstructured data into a structured triple form for subsequent data analysis and processing. First, the server performs natural language processing on the unstructured data to extract key information. Unstructured data usually exists in the form of natural language text, such as news reports, legal documents, patent specifications, etc. The server uses natural language processing technology to analyze and understand these texts. Specifically, the server first performs word segmentation and part-of-speech tagging on the text, splitting the text into individual words and tagging the part of speech of each word, such as nouns, verbs, adjectives, etc. Then, the server performs named entity recognition to identify named entities such as person names, place names, and organization names that appear in the text. Next, the server performs keyword extraction, extracting keywords in the text according to a predefined keyword dictionary or using algorithms such as TF-IDF, such as "company", "investment", "litigation", etc. Finally, the server performs syntactic analysis to identify the subject, predicate, and object sentence components in the text.
[0047] After extracting the key information, the server converts the key information into triples according to a preset template. The preset template defines the structure and format of the triples, usually including three parts: the subject, the predicate, and the object. The subject represents the doer of the action, the predicate represents the action or relationship, and the object represents the recipient or object of the action. The server maps the extracted key information into the preset template to generate the corresponding triples. For example, from the news text "Company A acquired 80% of the equity of Company B", the server extracts the key information "Company A", "acquired", "Company B", "80%", "equity". According to the preset template "[Company 1] acquires [Company 2] [equity percentage] equity", the triples "<Company A, acquired, Company B>" and "<Company A, acquired equity percentage, 80%>" are generated. After generating the triples, the server uses the triples as data elements of the target data set. Each triple represents a fact or relationship in the unstructured data, and multiple triples together constitute a structured representation of the unstructured data.
[0048] After completing the preprocessing of the structured data and the triple conversion of the unstructured data, the server associates and merges the two parts of the data according to the primary key or foreign key to obtain a complete and uniformly structured target data set. Each record in the target data set consists of several fields or triples.
[0049] Step S104: Based on the target data set, construct a data association graph. The nodes in the data association graph represent the data elements in the target data set, and the edges in the data association graph represent the association relationships between the various data elements. The association relationships are obtained through similarity calculations between the data elements.
[0050] In step S104, based on the target data set, a data association graph is constructed, which specifically includes: taking each data element in the target data set as a node of the data association graph; calculating the similarity between a first data element and a second data element, where the first data element and the second data element are any two of the multiple data elements; if it is determined that the similarity is greater than or equal to a preset similarity threshold, an edge is established between the first data element and the second data element to obtain the data association graph.
[0051] Specifically, the server constructs a data association graph based on the target data set. The data association graph is a graph model used to represent the association relationships between data elements. In the data association graph, nodes represent data elements, and edges represent the association relationships between data elements. First, the server takes each data element in the target data set as a node of the data association graph. The target data set includes structured data and unstructured data. The data elements can be field values in structured data, such as enterprise names, transaction amounts, etc.; or they can be triples obtained by converting unstructured data, such as "<Company A, investment, Company B>". The server creates a node for each data element, and the attributes of the node include the type, value, etc. of the data element. Then, the server calculates the similarity between data elements to determine whether there is an association relationship between them.
[0052] The server calculates the similarity for each pair of data elements in the target data set. If the similarity is greater than or equal to the preset similarity threshold, it is considered that there is an association relationship between these two data elements, and an edge is established between them. The setting of the preset similarity threshold needs to be determined according to specific business scenarios and data characteristics, and the optimal value can be found through experiments and tuning. This application does not limit this. Through the above steps, the server finally obtains a complete data association graph.
[0053] In a possible implementation manner, calculating the similarity between the first data element and the second data element specifically includes: extracting a first feature vector of the first data element and extracting a second feature vector of the second data element; calculating the cosine similarity between the first feature vector and the second feature vector, and taking the cosine similarity as the similarity between the first data element and the second data element.
[0054] Specifically, the server extracts the feature vectors of the first data element and the second data element. A feature vector is a mathematical representation used to describe the features and attributes of a data element. The method for extracting the feature vector depends on the type and structure of the data element: for text-type data elements such as enterprise names, news titles, etc., the server uses the bag-of-words model or the TF-IDF method to convert the text into a numerical vector of a fixed length. For example, for a news text, the server first tokenizes it to obtain a list of words; then, it counts the frequency of each word in the text to form a word frequency vector; finally, it normalizes the word frequency vector to obtain the feature vector of the news text.
[0055] For numerical-type data elements such as transaction amounts, stock prices, etc., the server directly uses their values as a component of the feature vector. For example, for a transaction record, the server extracts its numerical fields such as transaction amount, transaction time, transaction object, etc., to form a multi-dimensional numerical vector as the feature vector of the transaction.
[0056] After extracting the first feature vector of the first data element and the second feature vector of the second data element, the server calculates the cosine similarity between the two feature vectors. The cosine similarity can measure the cosine value of the angle between two vectors and can reflect the similarity of the vector directions. Finally, the server uses the calculated cosine similarity as the similarity between the first data element and the second data element. The server can set a preset similarity threshold. When the similarity between two data elements is greater than or equal to this threshold, it is considered that there is an association relationship between them, and an edge is established between them.
[0057] Step S105: In the data association graph, identify the region with abnormal association density. The region with abnormal association density is the region where the association density is abnormal, and the association density refers to the number of edges in a unit area.
[0058] In step S105, in the data association graph, identify the regions with abnormal association density, specifically including: performing grid processing on the data association graph, dividing the data association graph into multiple grid cells of the same size, where each grid cell includes at least one node; calculating the association density of each grid cell, where the association density is the actual number of edges within the grid cell divided by a preset number, and the preset number is the maximum number of edges that may exist within the grid cell; calculating the neighborhood association density of the first grid cell according to the association density, where the neighborhood association density is the average of the association densities of the second grid cells adjacent to the first grid cell, and the first grid cell is any one of the multiple grid cells; calculating the absolute value of the difference between the association density of the first grid cell and the neighborhood association density of the first grid cell, and taking the absolute value as the density difference degree of the first grid cell; determining the first grid cells with a density difference degree greater than or equal to a preset difference degree threshold as abnormal grid cells, and merging the adjacent abnormal grid cells into regions with abnormal association density.
[0059] Specifically, the server identifies the regions with abnormal association density in the data association graph. The regions with abnormal association density refer to the regions in the data association graph where the association density (i.e., the number of edges per unit area) is significantly higher or lower compared to the surrounding regions. These regions usually indicate that the association relationships between data elements are abnormally tight or sparse, and may reflect certain abnormal events or special patterns.
[0060] First, the server performs grid processing on the data association graph, dividing it into multiple grid cells of the same size. In the embodiments of the present application, grid processing can be understood as a spatial indexing method, which can transform an irregular graph structure into a regular grid structure. The server sets an appropriate grid size according to the size and density of the data association graph to balance the calculation efficiency and accuracy. Each grid cell contains at least one node in the data association graph and the edges connected to these nodes. Then, the server calculates the association density of each grid cell. The calculation formula for the association density is: Association density = Actual number of edges within the grid cell / Preset number. Among them, the preset number is the maximum number of edges that may exist within the grid cell, and can be calculated according to the combination number of the number of nodes within the grid cell. For example, for a grid cell containing n nodes, its preset number is n(n - 1) / 2. By dividing by the preset number, the association density can be normalized to the interval [0, 1], which is convenient for comparison between different grid cells.
[0061] Next, the server calculates the domain association density of each grid cell. The domain association density reflects the average association density level of the area around a grid cell. For the first grid cell, its domain association density is the average of the association densities of the second grid cells adjacent to it. Adjacent grid cells are the grid cells that are adjacent to the first grid cell in the up, down, left, right, and diagonal directions. By calculating the domain association density, the server can determine the local association density environment in which each grid cell is located.
[0062] After calculating the association density and the domain association density, the server calculates the density difference degree of each grid cell. The calculation formula for the density difference degree is: density difference degree = |association density of the first grid cell - domain association density of the first grid cell|. The density difference degree measures the degree of difference between the association density of a grid cell and the association density of its surrounding area. The larger the difference degree, the more likely the grid cell is an abnormal area.
[0063] Finally, the server identifies the abnormal grid cells according to a preset difference degree threshold, and merges the adjacent abnormal grid cells into an association density abnormal area. The server marks the grid cells with a density difference degree greater than or equal to the preset difference degree threshold as abnormal grid cells. The setting of the preset difference degree threshold is determined according to the distribution characteristics of the data and business requirements. It should be able to detect real abnormal areas while minimizing false alarms. The specific value of the preset difference degree threshold is not limited in this application. The server uses a region growing algorithm, with the abnormal grid cells as seeds, continuously absorbing the surrounding abnormal grid cells until no new abnormal grid cells are added, forming a complete abnormal area.
[0064] For example, the server performs abnormal area detection on a data association graph of an enterprise. The data association graph contains 1000 nodes and 5000 edges, representing the association relationships between 1000 enterprises. The server divides the data association graph into 100 grid cells of the same size, with each cell containing 10 nodes. Then, the server calculates the association density and the domain association density of each cell, and finds that the density difference degrees of 5 cells exceed the preset threshold of 0.5. Next, the server merges these 5 abnormal cells, and finally obtains 2 association density abnormal areas. Step S106: Mark the data elements included in the association density abnormal area as abnormal data.
[0065] Step S106: Mark the data elements included in the association density abnormal area as abnormal data.
[0066] In step S106, the server marks the data elements included in the association density abnormal area as abnormal data. Abnormal data refers to the data elements corresponding to the nodes within the association density abnormal area.
[0067] Specifically, the server first identifies the nodes included in each associated density anomaly region. Since the associated density anomaly region is formed by merging multiple adjacent abnormal grid cells, the nodes within the anomaly region are the set of all nodes within these abnormal grid cells. The server traverses the grid cells within the anomaly region to obtain the list of nodes therein, forming a set of abnormal nodes.
[0068] Then, the server marks the data elements corresponding to each node in the set of abnormal nodes as abnormal data. The server looks up the data elements corresponding to the nodes in the original dataset based on the attribute information of the nodes, such as node ID, node type, etc., and marks them as abnormal. The marking method can be to attach a special label or field to the data element indicating that it belongs to the anomaly region; or to store the abnormal data elements separately in a dedicated abnormal dataset for separate management from the normal data.
[0069] Step S107: Determine the corresponding abnormal dataset according to the position information of the abnormal data in the original dataset.
[0070] In step S107, determining the corresponding abnormal dataset according to the position information of the abnormal data in the original dataset specifically includes: obtaining the position information of the abnormal data in the target dataset, where the position information includes the row number and column number of the abnormal data in the target dataset; according to the position information of the abnormal data in the target dataset, obtaining the position information of the abnormal data in the structured data and unstructured data; according to the position information of the abnormal data in the structured data, determining the target structured data associated with the abnormal data, and according to the position information of the abnormal data in the unstructured data, determining the target unstructured data associated with the abnormal data; merging the target structured data and the target unstructured data to obtain a subset of the original data; using the subset of the original data as the abnormal dataset corresponding to the abnormal data.
[0071] Specifically, the server determines the abnormal data set associated with the abnormal data according to the location information of the abnormal data in the original data set. The abnormal data set refers to the data set in the same triple as the abnormal data. Specifically, the server first obtains the location information of the abnormal data in the target data set. The target data set is formed by merging structured data and unstructured data, and the abnormal data has clear row numbers and column numbers in it. The row number represents the number of the record or document where the abnormal data is located, and the column number represents the specific location or field of the abnormal data in the record or document. The server searches for the location information of the abnormal data in the index structure of the target data set through the unique identifier of the abnormal data, such as the node ID, data element ID, etc. Then, the server reversely searches for the location information of the abnormal data in the original structured data and unstructured data according to the location information of the abnormal data in the target data set. Since the target data set is obtained by preprocessing and transformation of structured data and unstructured data, there is a certain mapping relationship between the location information of the abnormal data in the target data set and its location information in the original data. The server can accurately locate the abnormal data in the structured data and unstructured data through this mapping relationship.
[0072] For structured data, the server searches for the corresponding records and fields in the structured data table or database according to the row number and column number of the abnormal data, and extracts other field data in the same record as the abnormal data to form the target structured data. For unstructured data, the server searches for the corresponding document or text segment in the unstructured data document or corpus according to the row number of the abnormal data, and extracts other text data in the same document as the abnormal data to form the target unstructured data.
[0073] Finally, the server merges the extracted target structured data and target unstructured data to obtain a subset of the original data as the abnormal data set corresponding to the abnormal data. The abnormal data set inherits the attribute structure and content characteristics of the original data, but only contains the records and documents associated with the abnormal data, with a smaller scale and stronger pertinence.
[0074] Refer to Figure 2, this application also provides a device for identifying associated abnormal data sets based on big data. The device is a server, and the server includes an acquisition module 201 and a processing module 202; the acquisition module 201 is used to acquire a target enterprise and upstream and downstream enterprises associated with the target enterprise; the processing module 202 is used to construct an original data set according to the target enterprise, upstream enterprises and downstream enterprises. The original data set includes structured data and unstructured data; the processing module 202 is also used to preprocess the structured data, perform triple conversion on the unstructured data, and merge the preprocessed structured data and the converted unstructured data to obtain a target data set; the processing module 202 is also used to construct a data association graph based on the target data set. The nodes in the data association graph represent data elements in the target data set, and the edges in the data association graph represent the association relationships between the data elements. The association relationships are obtained by calculating the similarity between the data elements; the processing module 202 is also used to identify an associated density abnormal area in the data association graph. The associated density abnormal area is an area with abnormal associated density. The associated density refers to the number of edges in a unit area; the processing module 202 is also used to mark the data elements included in the associated density abnormal area as abnormal data; the processing module 202 is also used to determine the corresponding abnormal data set according to the position information of the abnormal data in the original data set.
[0075] In a possible implementation manner, when the processing module 202 identifies an associated density abnormal area in the data association graph, it specifically includes: the processing module 202 performs grid processing on the data association graph, divides the data association graph into multiple grid cells of the same size, and each grid cell includes at least one node; the processing module 202 calculates the associated density of each grid cell. The associated density is the actual number of edges in the grid cell divided by a preset number, and the preset number is the maximum number of edges that may exist in the grid cell; the processing module 202 calculates the domain associated density of the first grid cell according to the associated density. The domain associated density is the average value of the associated densities of the second grid cells adjacent to the first grid cell. The first grid cell is any one of the multiple grid cells; the processing module 202 calculates the absolute value of the difference between the associated density of the first grid cell and the domain associated density of the first grid cell, and uses the absolute value as the density difference degree of the first grid cell; the processing module 202 determines that the first grid cell with a density difference degree greater than or equal to a preset difference degree threshold is an abnormal grid cell, and merges adjacent abnormal grid cells into an associated density abnormal area.
[0076] In a possible implementation, the processing module 202 constructs an original data set according to the target enterprise, upstream enterprises, and downstream enterprises, specifically including: the acquisition module 201 acquires the structured data of the target enterprise, upstream enterprises, and downstream enterprises, and the structured data includes enterprise basic information data, financial data, and transaction data; the acquisition module 201 acquires the unstructured data of the target enterprise, upstream enterprises, and downstream enterprises, and the unstructured data includes enterprise news data, judicial litigation data, and patent data; the processing module 202 performs data cleaning and data integration on the structured data and unstructured data to obtain the original data set.
[0077] In a possible implementation, the steps for the processing module 202 to perform triple transformation on the unstructured data specifically include: the processing module 202 performs natural language processing on the unstructured data to extract key information; the processing module 202 converts the key information into triples according to a preset template, and the triples include a subject, a predicate, and an object; the processing module 202 uses the triples as data elements of the target data set.
[0078] In a possible implementation, the processing module 202 constructs a data association graph based on the target data set, specifically including: the processing module 202 uses each data element in the target data set as a node of the data association graph; the processing module 202 calculates the similarity between a first data element and a second data element, and the first data element and the second data element are any two of the multiple data elements; if the processing module 202 determines that the similarity is greater than or equal to a preset similarity threshold, an edge is established between the first data element and the second data element to obtain the data association graph.
[0079] In a possible implementation, the processing module 202 calculates the similarity between the first data element and the second data element, specifically including: the processing module 202 extracts a first feature vector of the first data element and extracts a second feature vector of the second data element; the processing module 202 calculates the cosine similarity between the first feature vector and the second feature vector and uses the cosine similarity as the similarity between the first data element and the second data element.
[0080] In a possible implementation, the processing module 202 determines the corresponding abnormal data set according to the location information of the abnormal data in the original data set, specifically including: the acquisition module 201 acquires the location information of the abnormal data in the target data set, and the location information includes the row number and column number of the abnormal data in the target data set; the processing module 202 acquires the location information of the abnormal data in the structured data and unstructured data according to the location information of the abnormal data in the target data set; the processing module 202 determines the target structured data associated with the abnormal data according to the location information of the abnormal data in the structured data, and determines the target unstructured data associated with the abnormal data according to the location information of the abnormal data in the unstructured data; the processing module 202 merges the target structured data and the target unstructured data to obtain a subset of the original data; and takes the subset of the original data as the abnormal data set corresponding to the abnormal data.
[0081] It should be noted that: when the device provided in the above embodiment realizes its functions, only the division of the above function modules is used for illustration. In practical applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be elaborated here.
[0082] This application also provides an electronic device. Refer to Figure 3 , Figure 3 FIG. is a schematic structural diagram of an electronic device provided in an embodiment of the present application. The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.
[0083] Among them, the communication bus 302 is used to realize the connection and communication between these components.
[0084] Among them, the user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.
[0085] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0086] Among them, the processor 301 may include one or more processing cores. The processor 301 uses various interfaces and circuits to connect various parts within the entire server. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by invoking data stored in the memory 305, it performs various functions of the server and processes data. Optionally, the processor 301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 301 may integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately by a single chip.
[0087] Among them, the memory 305 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store the data involved in the above-mentioned method embodiments. Optionally, the memory 305 may further be at least one storage device located far from the aforementioned processor 301. Refer to Figure 3 , the memory 305, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for the method of identifying an associated abnormal data set based on big data.
[0088] In Figure 3In the electronic device 300 shown, the user interface 303 is mainly used to provide an interface for the user to input and obtain the data input by the user; and the processor 301 can be used to call the application program stored in the memory 305 for the method of identifying the associated exception data set based on big data. When executed by one or more processors 301, the electronic device 300 is caused to execute one or more of the methods as described in the above embodiments. It should be noted that for the foregoing method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0089] The present application also provides a computer-readable storage medium storing instructions. When executed by one or more processors 301, the electronic device 300 is caused to execute one or more of the methods as described in the above embodiments.
[0090] In the above embodiments, the descriptions of the various embodiments each have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0091] In several implementation manners provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some service interfaces. The indirect couplings or communication connections of the devices or units can be in electrical or other forms.
[0092] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0093] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0094] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned memory includes: various media such as USB flash drives, mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0095] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, all equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Those skilled in the art will readily think of other implementation schemes of the present disclosure after considering the specification and the disclosure of the practical truth.
[0096] The present application aims to cover any variations, uses, or adaptive changes of the present disclosure. These variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A method for identifying associated abnormal data sets based on big data, characterized in that The method includes: Obtaining a target enterprise, as well as upstream and downstream enterprises associated with the target enterprise; Constructing an original data set according to the target enterprise, the upstream enterprise, and the downstream enterprise, where the original data set includes structured data and unstructured data; Preprocessing the structured data, performing triple transformation on the unstructured data, and combining the preprocessed structured data and the transformed unstructured data to obtain a target data set; Based on the target data set, constructing a data association graph, where the nodes in the data association graph represent data elements in the target data set, the edges in the data association graph represent the association relationships between the respective data elements, and the association relationships are obtained by calculating the similarity between the data elements; In the data association graph, identifying an area with abnormal association density, where the area with abnormal association density is an area where the association density is abnormal, and the association density refers to the number of edges per unit area; Marking the data elements included in the area with abnormal association density as abnormal data; Determining a corresponding abnormal data set according to the position information of the abnormal data in the original data set; The step of identifying an area with abnormal association density in the data association graph specifically includes: Performing grid processing on the data association graph, dividing the data association graph into multiple grid units of the same size, and each grid unit includes at least one node; Calculating the association density of each grid unit, where the association density is the actual number of edges in the grid unit divided by a preset number, and the preset number is the maximum number of edges in the grid unit; According to the association density, calculating the neighborhood association density of the first grid unit, where the neighborhood association density is the average of the association densities of the second grid units adjacent to the first grid unit, and the first grid unit is any one of the multiple grid units; Calculating the absolute value of the difference between the association density of the first grid unit and the neighborhood association density of the first grid unit, and taking the absolute value as the density difference degree of the first grid unit; Determining the first grid unit with a density difference degree greater than or equal to a preset difference degree threshold as an abnormal grid unit, and merging adjacent abnormal grid units into the area with abnormal association density.
2. The method according to claim 1, wherein The step of constructing an original data set according to the target enterprise, the upstream enterprise, and the downstream enterprise specifically includes: Obtaining the structured data of the target enterprise, the upstream enterprise, and the downstream enterprise, where the structured data includes enterprise basic information data, financial data, and transaction data; Obtaining the unstructured data of the target enterprise, the upstream enterprise, and the downstream enterprise, where the unstructured data includes enterprise news data, judicial litigation data, and patent data; Performing data cleaning and data integration on the structured data and the unstructured data to obtain the original data set.
3. The method according to claim 1, wherein The step of performing triple transformation on the unstructured data specifically includes: Performing natural language processing on the unstructured data to extract key information; Convert the key information into triples according to a preset template, where the triples include a subject, a predicate, and an object; Use the triples as data elements of the target data set.
4. The method according to claim 1, wherein Based on the target data set, construct a data association graph, specifically including: Use each data element in the target data set as a node of the data association graph; Calculate the similarity between a first data element and a second data element, where the first data element and the second data element are any two of the multiple data elements; If it is determined that the similarity is greater than or equal to a preset similarity threshold, establish an edge between the first data element and the second data element to obtain the data association graph.
5. The method according to claim 4, wherein The calculation of the similarity between the first data element and the second data element specifically includes: Extract a first feature vector of the first data element and extract a second feature vector of the second data element; Calculate the cosine similarity between the first feature vector and the second feature vector, and use the cosine similarity as the similarity between the first data element and the second data element.
6. The method according to claim 1, wherein The determination of the corresponding abnormal data set according to the position information of the abnormal data in the original data set specifically includes: Obtain the position information of the abnormal data in the target data set, where the position information includes the row number and column number of the abnormal data in the target data set; According to the position information of the abnormal data in the target data set, obtain the position information of the abnormal data in the structured data and the unstructured data; According to the position information of the abnormal data in the structured data, determine the target structured data associated with the abnormal data, and according to the position information of the abnormal data in the unstructured data, determine the target unstructured data associated with the abnormal data; Merge the target structured data and the target unstructured data to obtain a subset of the original data; Use the subset of the original data as the abnormal data set corresponding to the abnormal data.
7. An apparatus for identifying associated abnormal data sets based on big data, characterized in that, The device is used to execute the method according to any one of claims 1-6. The device includes an acquisition module (201) and a processing module (202), where: The acquisition module (201) is used to acquire a target enterprise and upstream and downstream enterprises associated with the target enterprise; The processing module (202) is used to construct an original data set according to the target enterprise, the upstream enterprise, and the downstream enterprise, where the original data set includes structured data and unstructured data; The processing module (202) is further used to preprocess the structured data, perform triple conversion on the unstructured data, and merge the preprocessed structured data and the converted unstructured data to obtain a target data set; The processing module (202) is further configured to construct a data association graph based on the target data set. Nodes in the data association graph represent data elements in the target data set, and edges in the data association graph represent the association relationships between the respective data elements. The association relationships are obtained by calculating the similarity between the data elements; The processing module (202) is further configured to identify an association density abnormal region in the data association graph. The association density abnormal region is a region with abnormal association density. The association density refers to the number of edges in a unit region; The processing module (202) is further configured to mark the data elements included in the association density abnormal region as abnormal data; The processing module (202) is further configured to determine a corresponding abnormal data set according to the position information of the abnormal data in the original data set; The processing module (202) is further configured to perform grid processing on the data association graph, divide the data association graph into a plurality of grid cells of the same size, and each grid cell includes at least one node; calculate the association density of each grid cell, where the association density is the actual number of edges in the grid cell divided by a preset number, and the preset number is the maximum number of edges in the grid cell; calculate the domain association density of the first grid cell according to the association density, where the domain association density is the average value of the association densities of the second grid cells adjacent to the first grid cell, and the first grid cell is any one of the plurality of grid cells; calculate the absolute value of the difference between the association density of the first grid cell and the domain association density of the first grid cell, and use the absolute value as the density difference degree of the first grid cell; determine that the first grid cell with the density difference degree greater than or equal to a preset difference degree threshold is an abnormal grid cell, and merge adjacent abnormal grid cells into the association density abnormal region.
8. An electronic device, characterized in that, It includes a processor (301), a memory (305), a user interface (303), and a network interface (304). The memory (305) is used to store instructions. The user interface (303) and the network interface (304) are used to communicate with other devices. The processor (301) is used to execute the instructions stored in the memory (305) so that the electronic device (300) executes the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1-6 is executed.
Citation Information
Patent Citations
Internet financial risk monitoring system based on knowledge map
CN109064318A
Safety monitoring method and device and storage medium
CN117353954A