Abnormal data source identification method and device, equipment, medium and product

By constructing the knowledge graph of the target data and performing hierarchical analysis based on abnormal information, the problem of low efficiency and accuracy of abnormal data source identification in big data scenarios is solved, and efficient and accurate automatic identification of abnormal data sources is achieved.

CN120105288APending Publication Date: 2025-06-06BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510120681.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In big data scenarios, business data comes from data sources from multiple dimensions, making it difficult to efficiently and accurately detect abnormal data sources when data abnormalities occur.

Method used

By obtaining the knowledge graph corresponding to the target data, the knowledge graph includes multiple entities, each entity corresponds one by one to multiple data sources, and is divided into multiple levels based on the inclusion relationship. Based on the exception information of the target data, the exception entity is obtained in the current entity at the current level in multiple levels, and the exception data source is determined in multiple data sources.

Benefits of technology

Automatic identification of abnormal data sources is realized. Compared with manual analysis, abnormal data sources can be identified efficiently and accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105288A_ABST
    Figure CN120105288A_ABST
Patent Text Reader

Abstract

The invention provides an abnormal data source identification method and device, equipment, a medium and a product, and relates to the technical field of artificial intelligence, in particular to the technical fields of big data, cloud computing and the like. The abnormal data source identification method comprises: obtaining a knowledge graph corresponding to target data, the knowledge graph comprising a plurality of entities, the plurality of entities being in one-to-one correspondence with a plurality of data sources of the target data, and the plurality of entities being divided into a plurality of levels based on an inclusion relationship; on the basis of the abnormal information of the target data, obtaining an abnormal entity in the current entity of the current hierarchy in the plurality of hierarchies; and determining an abnormal data source in the plurality of data sources based on the abnormal entity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the technical fields of big data, cloud computing, etc., and in particular to a method, device, equipment, medium and product for identifying an abnormal data source. Background Art

[0002] In big data scenarios, business data usually comes from data sources of multiple dimensions. When business data is abnormal, it is necessary to detect the abnormal data source efficiently and accurately. Summary of the invention

[0003] The present disclosure provides a method, apparatus, device, medium and product for identifying an abnormal data source.

[0004] According to one aspect of the present disclosure, a method for identifying an abnormal data source is provided, comprising: obtaining a knowledge graph corresponding to target data, the knowledge graph comprising a plurality of entities, the plurality of entities corresponding one-to-one to a plurality of data sources of the target data, and the plurality of entities being divided into a plurality of levels based on an inclusion relationship; based on abnormal information of the target data, obtaining an abnormal entity in a current entity of a current level among the plurality of levels; and determining an abnormal data source among the plurality of data sources based on the abnormal entity.

[0005] According to another aspect of the present disclosure, there is provided an abnormal data source identification device, including: a first acquisition module, used to acquire a knowledge graph corresponding to target data, the knowledge graph including multiple entities, the multiple entities corresponding one-to-one to multiple data sources of the target data, and the multiple entities are divided into multiple levels based on an inclusion relationship; a second acquisition module, used to acquire abnormal entities in a current entity of a current level among the multiple levels based on abnormal information of the target data; and a determination module, used to determine an abnormal data source among the multiple data sources based on the abnormal entity.

[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any method as described in any of the above aspects.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any one of the methods according to any one of the above aspects.

[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the computer program implements any one of the methods described in any one of the above aspects.

[0009] According to the embodiments of the present disclosure, abnormal data sources can be identified efficiently and accurately.

[0010] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0012] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;

[0013] Figure 2 is a schematic diagram of an implementation system for implementing an embodiment of the present disclosure;

[0014] Figure 3 is a schematic diagram of the overall architecture for implementing the embodiments of the present disclosure;

[0015] Figure 4 is a schematic diagram according to a second embodiment of the present disclosure;

[0016] Figure 5 is a schematic diagram according to a third embodiment of the present disclosure;

[0017] Figure 6 It is a schematic diagram of an electronic device used to implement the abnormal data source identification method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0018] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0019] Taking the advertising system as an example, advertising revenue includes multiple dimensions of revenue. When the advertising revenue is abnormal, it is necessary to detect the dimension where the abnormality occurs.

[0020] In related technologies, manual methods are usually used to detect abnormal dimensions, but there are problems with efficiency and accuracy.

[0021] In order to efficiently and accurately detect abnormal data sources, the present disclosure provides the following embodiments.

[0022] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure. This embodiment provides a method for identifying an abnormal data source, such as Figure 1 As shown, the method includes:

[0023] 101. Obtain a knowledge graph corresponding to the target data, wherein the knowledge graph includes multiple entities, the multiple entities correspond one-to-one to multiple data sources of the target data, and the multiple entities are divided into multiple levels based on the inclusion relationship between the entities.

[0024] 102. Based on the abnormal information of the target data, determine an abnormal entity in a current entity of a current level among the multiple levels.

[0025] 103. Determine an abnormal data source from the candidate data sources based on the abnormality detection result of the current entity.

[0026] The target data is the data to be analyzed, and the target data comes from multiple data sources.

[0027] For example, in an advertising system, the target data refers to advertising revenue, and multiple data sources refer to multiple revenue dimensions.

[0028] In a big data scenario, there is usually an inclusion relationship between multiple data sources. For example, advertising revenue includes a first income dimension and a second income dimension, and the first income dimension further includes a third income dimension and a fourth income dimension.

[0029] For multiple data sources of target data, a knowledge graph can be constructed based on these data sources and their inclusion relationships. The knowledge graph includes entities and relationships. Each entity corresponds to a data source, and the relationship is specifically an inclusion relationship.

[0030] Specifically, the knowledge graph can be represented by nodes and edges. Each node represents an entity, and each entity corresponds to a data source. Edges are directed edges, which are used to represent inclusion relationships. For example, based on the above example, the knowledge graph contains four entities, corresponding to the four income dimensions mentioned above, and assuming that the four income dimensions correspond to the first entity to the fourth entity, there is an edge from the first entity to the third entity, and an edge from the first entity to the fourth entity.

[0031] Based on the inclusion relationship, multiple entities are divided into multiple levels, such as the first entity mentioned above is the upper level of the third entity and the fourth entity.

[0032] When an abnormality occurs in the target data, the abnormality information of the target data can be obtained. The abnormality information includes, for example, abnormal time, abnormal type, etc. The abnormality type includes, for example, week-on-week comparison, day-on-day comparison, etc.

[0033] After obtaining the abnormal information, it can be analyzed layer by layer based on the knowledge graph to locate the abnormal data source.

[0034] Specifically, a level currently being processed among multiple levels may be referred to as a current level, an entity in the current level may be referred to as a current entity, and the current entity may be one or more.

[0035] For the current level, based on the abnormal information of the target data, the abnormal entity is determined in the current entity, and the abnormal data source is determined based on the abnormal entity.

[0036] For example, the current level is the first level, and the first level includes a first entity and a second entity. When an abnormality occurs in the target data, the abnormal entity is determined in the first entity and the second entity based on the abnormal information. If the abnormal entity is the second entity, then the abnormal data source is determined based on the second entity, such as taking the second income dimension corresponding to the second entity as the abnormal data source.

[0037] In this embodiment, when an abnormality occurs in the target data, the abnormal data source is determined based on the knowledge graph corresponding to the target data, so that automatic identification of the abnormal data source can be achieved. Compared with manual analysis methods, the abnormal data source can be identified efficiently and accurately.

[0038] In order to better understand the present disclosure, the application scenarios involved in the present disclosure are described as follows:

[0039] Figure 2 It is a schematic diagram of an implementation system for implementing the embodiments of the present disclosure.

[0040] like Figure 2 As shown, this scenario involves: a monitoring system 201, a positioning system 202 and a database 203.

[0041] The monitoring system 201 is used to monitor whether the target data is abnormal, and when the target data is abnormal, an abnormal event is generated to the positioning system 202. The abnormal event includes abnormal information, such as abnormal time, abnormal type, etc.

[0042] The positioning system 202 is used to identify the abnormal data source after receiving the abnormal event, so as to determine the abnormal data source among multiple data sources of the target data.

[0043] Specifically, the positioning system 202 may perform identification based on a pre-built knowledge graph.

[0044] Database 203 is used to store the knowledge graph corresponding to the target data.

[0045] The knowledge graph is constructed based on multiple data sources of target data and their inclusion relationships.

[0046] The knowledge graph includes multiple entities, which correspond one-to-one to multiple data sources, and the multiple entities are divided into multiple levels based on the inclusion relationship.

[0047] Figure 3 It is a schematic diagram of the knowledge graph and entity hierarchy provided according to an embodiment of the present disclosure.

[0048] like Figure 3 As shown, assuming that there are 6 data sources of the target data, 6 entities are created and represented by the first entity to the sixth entity respectively; assuming that the second data source (corresponding to the second entity) contains the fourth data source (corresponding to the fourth entity), the fourth data source contains the sixth data source (corresponding to the sixth entity), and the third data source (corresponding to the third entity) contains the fifth data source (corresponding to the fifth entity), then an edge from the second entity to the fourth entity, an edge from the fourth entity to the sixth entity, and an edge from the third entity to the fifth entity are constructed.

[0049] In addition, attributes can be created for each entity. Attributes include monitoring indicators, which are used to obtain data from the data source corresponding to the entity. Taking the first entity as an example, its attributes include the first monitoring indicator. Based on the first monitoring indicator, relevant data corresponding to the abnormal information can be obtained from the first data source corresponding to the first entity, such as the week-on-week data corresponding to the abnormal time. The remaining entities are similar and also have monitoring indicators to obtain relevant data based on the monitoring indicators. For simplicity, Figure 3 Monitoring indicators for other entities are not shown.

[0050] based on Figure 3 The knowledge graph of a network can be divided into multiple levels, such as the first level, the second level and the third level from top to bottom. The first level includes the first entity, the second entity and the third entity, the second level includes the fourth entity and the fifth entity, and the third level includes the sixth entity. Each level can include one or more entities. The first level is the top level.

[0051] When identifying abnormal data sources based on knowledge graphs, initially, the top level (such as Figure 3The first level in the example is used as the current level, and based on the relevant data corresponding to the abnormal information of the current entities (first entity to third entity) in the current level, whether the corresponding entity is an abnormal entity is analyzed based on preset rules. For example, based on the first monitoring indicator, relevant data corresponding to the abnormal information is obtained from the first data source, and the relevant data is analyzed. If the relevant data does not meet the preset conditions, such as the relevant data is the week-on-week data of the first data source, if the week-on-week data is not within the preset range, it is determined that the first entity is an abnormal entity.

[0052] Multiple current entities at the current level may be analyzed in parallel, such as analyzing in parallel whether the first entity, the second entity, and the third entity are abnormal entities.

[0053] After the abnormal entity is determined, if the abnormal entity has a lower-level entity, the above process is repeated with the lower-level entity as the new current entity until the abnormal entity with the smallest granularity, that is, the lowest-level level, is obtained.

[0054] Afterwards, the data sources corresponding to the abnormal entities with the smallest granularity are summarized to obtain the final abnormal data source. For example, the fifth entity and the sixth entity are abnormal entities. Since these two entities are the entities with the smallest granularity, the fifth data source corresponding to the fifth entity and the sixth data source corresponding to the sixth entity are used as the final abnormal data sources.

[0055] In this way, the final abnormal data source can be obtained through level by level.

[0056] In combination with the above application scenarios, the present disclosure also provides the following embodiments.

[0057] Figure 4 is a schematic diagram according to the second embodiment of the present disclosure. This embodiment provides a method for identifying an abnormal data source, such as Figure 4 As shown, the method includes:

[0058] 401. Obtain a knowledge graph corresponding to the target data, wherein the knowledge graph includes multiple entities, the multiple entities correspond one-to-one to multiple data sources of the target data, and the multiple entities are divided into multiple levels based on an inclusion relationship.

[0059] 402. Based on the abnormal information of the target data, determine an abnormal entity in a current entity of a current level among the multiple levels.

[0060] Initially, the topmost level is used as the current level. For example, see Figure 3 ,Initially, the first level is taken as the current level, and the entities in the first level are taken as the current entities.

[0061] Afterwards, relevant data corresponding to the abnormal information may be obtained from a data source corresponding to the current entity; and the relevant data may be analyzed to obtain the abnormal entity.

[0062] Specifically, a query statement of the knowledge graph can be generated based on the abnormal information, and the knowledge graph can be queried using the query statement to obtain relevant data, which is then analyzed. The abnormal entity is determined based on the analysis result. For example, if the relevant data of a current entity is not within a preset range, the current entity is determined as an abnormal entity.

[0063] In this embodiment, by acquiring and analyzing relevant data, abnormal entities can be acquired, and then abnormal data sources can be identified based on the abnormal entities, thereby improving processing accuracy.

[0064] Furthermore, if there are multiple current entities, the multiple current entities can be analyzed in parallel to determine the abnormal entity.

[0065] For example, see Figure 3 If the current level is the first level, the current entities include the first entity, the second entity and the third entity, then the three entities can be analyzed in parallel to determine the abnormal entity among the three entities.

[0066] In this embodiment, processing efficiency can be improved by processing multiple current entities in parallel.

[0067] 403. Determine whether the abnormal entity has a lower-level entity. If so, execute 404; otherwise, execute 405.

[0068] 404. Take the lower-layer entity as a new current entity, and then repeat 403 and subsequent steps.

[0069] For example, refer to Figure 3 If the fourth entity is an abnormal entity, since the fourth entity has a lower-level entity (the sixth entity), the sixth entity is used as the new current entity and the relevant process is re-executed.

[0070] 405. Use the data source corresponding to the abnormal entity as the abnormal data source.

[0071] For example, refer to Figure 3 If the fifth entity is an abnormal entity, since the fifth entity has no lower-level entity, the data source corresponding to the fifth entity is taken as the abnormal data source.

[0072] In this embodiment, multiple data sources can be integrated through the knowledge graph, and then relevant data of the abnormal information can be obtained based on the knowledge graph, and the abnormal data source can be located based on the relevant data. In this way, the abnormal data source can be located by querying the knowledge graph, which can improve processing efficiency and accuracy.

[0073] In this embodiment, when the abnormal entity has a lower-level entity, the lower-level entity is analyzed as a new current entity, so that entities at multiple levels can be analyzed layer by layer, thereby improving the comprehensiveness and accuracy of the processing.

[0074] In this embodiment, when the abnormal entity does not have a lower-level entity, the data source corresponding to the abnormal entity is used as the abnormal data source, so that the abnormal data source with the smallest granularity can be obtained, thereby improving the accuracy of locating the abnormal data source.

[0075] Figure 5 is a schematic diagram according to the third embodiment of the present disclosure, which provides an abnormal data source identification device. Figure 5 As shown, the device 500 includes: a first acquisition module 501, a second acquisition module 502 and a determination module 503.

[0076] The first acquisition module 501 is used to acquire the knowledge graph corresponding to the target data, wherein the knowledge graph includes multiple entities, and the multiple entities correspond one-to-one to multiple data sources of the target data, and the multiple entities are divided into multiple levels based on the inclusion relationship; the second acquisition module 502 is used to acquire abnormal entities in the current entity of the current level in the multiple levels based on the abnormal information of the target data; the determination module 503 is used to determine the abnormal data source in the multiple data sources based on the abnormal entity.

[0077] In this embodiment, when an abnormality occurs in the target data, the abnormal data source is determined based on the knowledge graph corresponding to the target data, so that automatic identification of the abnormal data source can be achieved. Compared with manual analysis methods, the abnormal data source can be identified efficiently and accurately.

[0078] In some embodiments, the determining module 503 is further configured to:

[0079] If the abnormal entity has no lower-layer entity, the data source corresponding to the abnormal entity is used as the abnormal data source.

[0080] In this embodiment, when the abnormal entity does not have a lower-level entity, the data source corresponding to the abnormal entity is used as the abnormal data source, so that the abnormal data source with the smallest granularity can be obtained, thereby improving the accuracy of locating the abnormal data source.

[0081] In some embodiments, the apparatus 500 further includes:

[0082] The third acquisition module is used to use the lower-layer entity as a new current entity if the abnormal entity has a lower-layer entity.

[0083] In this embodiment, when the abnormal entity has a lower-level entity, the lower-level entity is analyzed as a new current entity, so that entities at multiple levels can be analyzed layer by layer, thereby improving the comprehensiveness and accuracy of the processing.

[0084] In some embodiments, the second acquisition module 502 is further configured to:

[0085] In the data source corresponding to the current entity, obtain relevant data corresponding to the abnormal information;

[0086] The relevant data is analyzed to obtain the abnormal entity.

[0087] In this embodiment, by acquiring and analyzing relevant data, abnormal entities can be acquired, and then abnormal data sources can be identified based on the abnormal entities, thereby improving processing accuracy.

[0088] In some embodiments, the second acquisition module 502 is further configured to:

[0089] If there are multiple current entities, the current entities are processed in parallel based on the exception information to obtain the exception entity.

[0090] In this embodiment, processing efficiency can be improved by processing multiple current entities in parallel.

[0091] It can be understood that in the embodiments of the present disclosure, the same or similar contents in different embodiments can be referenced to each other.

[0092] It can be understood that the “first”, “second”, etc. in the embodiments of the present disclosure are only used for distinction and do not indicate the degree of importance, time sequence, etc.

[0093] It is understandable that unless there is any special limitation on the sequence of steps in the process, it means that the timing relationship between these steps is not limited.

[0094] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0095] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0096] Figure 6A schematic block diagram of an example electronic device 600 that can be used to implement an embodiment of the present disclosure is shown. The electronic device 600 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0097] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 606 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0098] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0099] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as the abnormal data source identification method. For example, in some embodiments, the abnormal data source identification method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the abnormal data source identification method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the abnormal data source identification method in any other appropriate manner (e.g., by means of firmware).

[0100] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0101] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable task processing device, so that the program code, when executed by the processor or controller, implements the functions / operations specified in the flow chart and / or block diagram. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0102] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0103] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0104] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0105] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.

[0106] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0107] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for identifying an abnormal data source, comprising: Obtain a knowledge graph corresponding to the target data, wherein the knowledge graph includes a plurality of entities, the plurality of entities correspond one-to-one to a plurality of data sources of the target data, and the plurality of entities are divided into a plurality of levels based on a containment relationship; Based on the abnormal information of the target data, obtaining an abnormal entity in a current entity of a current level among the multiple levels; Based on the abnormal entity, an abnormal data source is determined among the plurality of data sources.

2. The method according to claim 1, wherein: The determining, based on the abnormal entity, an abnormal data source from among the multiple data sources comprises: If the abnormal entity has no lower-layer entity, the data source corresponding to the abnormal entity is used as the abnormal data source.

3. The method according to claim 1, further comprising: If the abnormal entity has a lower-level entity, the lower-level entity is used as a new current entity.

4. The method according to claim 1, wherein: The acquiring, based on the abnormal information of the target data, an abnormal entity in a current entity of a current level among the multiple levels, comprises: In the data source corresponding to the current entity, obtain relevant data corresponding to the abnormal information; The relevant data is analyzed to obtain the abnormal entity.

5. The method according to claim 1, wherein: The acquiring, based on the abnormal information of the target data, an abnormal entity in a current entity of a current level among the multiple levels, comprises: If there are multiple current entities, the current entities are processed in parallel based on the exception information to obtain the exception entity.

6. An abnormal data source identification device, comprising: A first acquisition module is used to acquire a knowledge graph corresponding to the target data, wherein the knowledge graph includes a plurality of entities, the plurality of entities correspond one-to-one to a plurality of data sources of the target data, and the plurality of entities are divided into a plurality of levels based on an inclusion relationship; A second acquisition module, configured to acquire an abnormal entity from a current entity of a current level among the multiple levels based on the abnormal information of the target data; A determination module is used to determine an abnormal data source from among the multiple data sources based on the abnormal entity.

7. The device according to claim 6, wherein: The determination module is further used for: If the abnormal entity has no lower-layer entity, the data source corresponding to the abnormal entity is used as the abnormal data source.

8. The apparatus according to claim 6, further comprising: The third acquisition module is used to use the lower-layer entity as a new current entity if the abnormal entity has a lower-layer entity.

9. The device according to claim 6, wherein: The second acquisition module is further used for: In the data source corresponding to the current entity, obtain relevant data corresponding to the abnormal information; The relevant data is analyzed to obtain the abnormal entity.

10. The device according to claim 6, wherein: The second acquisition module is further used for: If there are multiple current entities, the current entities are processed in parallel based on the exception information to obtain the exception entity.

11. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5.

13. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 5.