Data processing method and device, electronic equipment and storage medium

CN118626534BActive Publication Date: 2026-09-22TENCENT TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310246358.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-07
Publication Date
2026-09-22
Estimated Expiration
2043-03-07

AI Technical Summary

Technical Problem

[0003]相关技术中,主要关注于已获得的对象画像的应用,而缺乏关注所获得的对象画像的质量

Benefits of technology

[0018]本申请提供了一种更具全局性的对象集群画像,该画像用于表征针对目标数据源的加工链路上各个任务阶段的目标阶段指标数据,同时目标阶段指标数据包括数据维度的指标数据和任务执行维度的指标数据。通过这样的画像可以进行针对加工链路的异常监测,进而实现准确有效的画像质量监测,从而为相关用户提供更具针对性的服务。这样的画像可以为画像的运维管理提供整体性数据,提高运维管理的准确度和效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118626534B_ABST
    Figure CN118626534B_ABST
Patent Text Reader

Abstract

The application relates to a data processing method and device, electronic equipment and a storage medium. The method comprises the following steps: in response to a received portrait query request, extracting a preset condition from the portrait query request, and determining a target data source meeting the preset condition; obtaining target portrait data of a target object cluster, and performing abnormal monitoring on a processing link based on the target portrait data, the target portrait data being used to represent target stage index data of each task stage on the processing link of the target data source. The application provides a more global object cluster portrait. Through the portrait, accurate and effective portrait quality monitoring can be realized, so that more targeted services can be provided for related users. The portrait can provide overall data for portrait operation and management, and improve the accuracy and efficiency of operation and management. The embodiments of the application can be applied to various scenes such as cloud technology, artificial intelligence, intelligent transportation and intelligent entertainment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet communication technology, and in particular to a data processing method, apparatus, electronic device and storage medium. Background Technology

[0002] With the development of internet communication technology, various internet products have emerged, providing users with service experiences. Object profiling refers to the technology used to describe the characteristics of an object. It involves acquiring the object's tag attributes and then using these attributes to describe various aspects of the object's characteristics. Object profiling can be used to uncover object needs and analyze object preferences, and by matching object profiles, more efficient and targeted information can be provided.

[0003] Current technologies primarily focus on the application of existing object profiles, neglecting the quality of those profiles. Therefore, there is a need to provide accurate and effective profile quality monitoring solutions. Summary of the Invention

[0004] To address at least one of the aforementioned technical problems, this application provides a data processing method, apparatus, electronic device, and storage medium:

[0005] According to a first aspect of this application, a data processing method is provided, the method comprising:

[0006] In response to a received profile query request, preset conditions are extracted from the profile query request, and a target data source that meets the preset conditions is determined. The target data source consists of operation data of the target object cluster.

[0007] Obtain target profile data of the target object cluster to perform anomaly monitoring on the processing link based on the target profile data. The target profile data is used to characterize target stage indicator data for each task stage on the processing link of the target data source. The target stage indicator data includes a first type of indicator data and a second type of indicator data. The first type of indicator data is used to describe the data flow of the task stage, and the second type of indicator data is used to describe the task execution of the task stage.

[0008] According to a second aspect of this application, a data processing method is provided, applied to a data processing system, the data processing system comprising a first type of server cluster and multiple second type of server clusters, the first type of server clusters being used to execute the data processing method as described in the first aspect, and the multiple second type of server clusters being respectively responsible for executing tasks at each task stage in the processing chain, the method comprising:

[0009] For each task stage, the second type of server cluster corresponding to the task stage generates stage indicator data to be reported using a preset time window as the time statistical caliber, and reports the stage indicator data to be reported to the first type of server cluster.

[0010] According to a third aspect of this application, a data processing apparatus is provided, the apparatus comprising:

[0011] Response module: Used to respond to a received profile query request, extract preset conditions from the profile query request, and determine the target data source that meets the preset conditions. The target data source consists of operation data of the target object cluster.

[0012] Acquisition Module: Used to acquire target profile data of the target object cluster, so as to perform anomaly monitoring on the processing link based on the target profile data. The target profile data is used to characterize the target stage indicator data of each task stage on the processing link for the target data source. The target stage indicator data includes a first type of indicator data and a second type of indicator data. The first type of indicator data is used to describe the data flow of the task stage, and the second type of indicator data is used to describe the task execution of the task stage.

[0013] According to a fourth aspect of this application, an electronic device is provided, the electronic device including at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores at least one instruction or at least one program, the at least one instruction or at least one program being loaded and executed by the at least one processor to implement the data processing method as described in the first aspect or the second aspect.

[0014] According to a fifth aspect of this application, a computer-readable storage medium is provided, wherein at least one instruction or at least one program is stored therein, the at least one instruction or at least one program being loaded and executed by a processor to implement the data processing method as described in the first or second aspect.

[0015] According to a sixth aspect of this application, a computer program product is provided, the computer program product comprising at least one instruction or at least one program segment, the at least one instruction or at least one program segment being loaded and executed by a processor to implement the data processing method as described in the first or second aspect.

[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application.

[0017] Implementing this application will have the following beneficial effects:

[0018] This application provides a more global object cluster profile, which characterizes the target stage indicator data for each task stage in the processing chain of a target data source. The target stage indicator data includes both data-dimensional and task execution-dimensional indicator data. This profile enables anomaly monitoring of the processing chain, thereby achieving accurate and effective profile quality monitoring and providing more targeted services to relevant users. Such a profile provides holistic data for profile operation and maintenance management, improving the accuracy and efficiency of operation and maintenance management.

[0019] Other features and aspects of this application will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0020] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This diagram illustrates an application environment according to an embodiment of the present application.

[0022] Figure 2 A flowchart illustrating a data processing method according to an embodiment of this application is shown;

[0023] Figure 3 This diagram illustrates a process for obtaining target profile data of a target object cluster according to an embodiment of this application.

[0024] Figure 4 This diagram illustrates a process for displaying target profile data according to an embodiment of this application.

[0025] Figure 5 A flowchart illustrating a data processing method according to an embodiment of this application is shown;

[0026] Figure 6 A schematic diagram of the interface of the back-end panel display dashboard according to an embodiment of this application is shown;

[0027] Figure 7 Also shown is a schematic diagram of the interface of the back-end panel display dashboard according to an embodiment of this application;

[0028] Figure 8 This diagram illustrates a device block diagram according to an embodiment of the present application;

[0029] Figure 9 A schematic diagram of an electronic device according to an embodiment of this application is shown. Detailed Implementation

[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0032] Various exemplary embodiments, features, and aspects of this application will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0033] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0034] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0035] Furthermore, to better illustrate this application, numerous specific details are provided in the following detailed description. Those skilled in the art should understand that this application can be implemented without certain specific details. In some instances, methods, means, components, and circuits well-known to those skilled in the art have not been described in detail in order to highlight the main points of this application.

[0036] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0037] QPS (Queries per second): Queries per second.

[0038] Spark: An open-source cluster computing environment that is a fast and general-purpose computing engine designed for large-scale data processing.

[0039] The crontab command: a type of computer program instruction.

[0040] Oozie: An open-source workflow and collaboration service engine.

[0041] Operational data store (ODS): It is a type of database that is often used as a temporary area in a data warehouse.

[0042] TDW: A Distributed Data Warehouse

[0043] ATTA: A log pipeline system.

[0044] Data Warehouse Detail (DWD): This layer mainly cleanses and integrates the ODS layer data synchronized from the business database into corresponding fact tables.

[0045] MySQL: It is a relational database management system.

[0046] Redis (Remote Dictionary Server): A remote dictionary service.

[0047] BDB (Berkeley DB): It is an open-source file database.

[0048] YottaDB: It is a real-time database.

[0049] KeeWiDB: It is a key-value database.

[0050] Please see Figure 1 , Figure 1The diagram illustrates an application environment according to an embodiment of this application. The application environment may include a client 10 and a server 20. The client 10 and server 20 can be directly or indirectly connected via wired or wireless communication. Users (e.g., staff) can send a profile query request to the server 20 via the client 10. In response to the received profile query request, the server 20 extracts preset conditions from the profile query request and determines a target data source that meets the preset conditions. The target data source consists of operation data of a target object cluster. Then, it obtains target profile data of the target object cluster to perform anomaly monitoring on the processing link based on the target profile data. The target profile data is used to characterize target stage indicator data for each task stage on the processing link of the target data source. The target stage indicator data includes a first type of indicator data and a second type of indicator data. The first type of indicator data describes the data flow of the task stage, and the second type of indicator data describes the task execution of the task stage. Of course, the server 20 can also execute the data processing scheme provided in this embodiment based on its own generated profile query request. It should be noted that... Figure 1 This is just one example.

[0051] Client 10 can be a physical device such as a smartphone, computer (e.g., desktop computer, tablet, laptop), augmented reality (AR) / virtual reality (VR) device, digital assistant, smart voice interaction device (e.g., smart speaker), smart wearable device, smart home appliance, in-vehicle terminal, etc., or software running on the physical device, such as a computer program. The operating system corresponding to Client 10 can be Android, iOS (a mobile operating system developed by Apple), Linux, Microsoft Windows, etc.

[0052] The server-side component 20 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server may include network communication units, processors, and memory, etc. The server-side component 20 can provide backend services to the corresponding client 10.

[0053] In practical applications, the data processing solutions provided in this application can utilize technologies related to Artificial Intelligence (AI) and cloud storage. Artificial intelligence is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, artificial intelligence is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.

[0054] Cloud storage is a new concept that has been extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as a storage system) refers to a storage system that uses cluster applications, grid technology and distributed storage file systems to bring together a large number of storage devices of various types in the network (storage devices are also called storage nodes) to work together through application software or application interfaces to provide data storage and business access functions to the outside world.

[0055] It should be noted that when this application embodiment is applied to specific products or technologies for target data sources that are related to user information, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0056] Figure 2 This diagram illustrates a flow chart of a data processing method according to an embodiment of this application, such as... Figure 2 As shown, the method includes:

[0057] S201: In response to the received profile query request, extract preset conditions from the profile query request and determine the target data source that meets the preset conditions, wherein the target data source consists of operation data of the target object cluster;

[0058] In this embodiment, in response to a received profile query request, the server extracts preset conditions from the profile query request and determines a target data source that meets the preset conditions. The profile query request can be triggered and generated by a user (e.g., a staff member) through an interactive interface provided by the client and sent to the server. Alternatively, the profile query request can be triggered and generated by the server itself based on preset trigger conditions. For example, the server loads a profile query strategy for a specified object cluster, which includes preset trigger conditions. If the current condition (e.g., the current time) meets the preset trigger conditions, a profile query request is generated.

[0059] An object can refer to a user, account, etc. An object cluster includes at least one object. An object cluster profile is a description of the multi-dimensional characteristics of an object cluster, and an object cluster profile can also be composed of multi-dimensional characteristics. A profile query request aims to query an object cluster profile; the request indicates an intention to query the profile of a specific object cluster, and it limits the target data source by carrying preset conditions. The server extracts the preset conditions from the profile query request and determines the target data source that meets the preset conditions. The target data source consists of the operation data of the target object cluster. It can be understood that preset conditions are used to limit the target data source, and the target data source comes from the operation data of the target object cluster. Therefore, preset conditions can also be used to limit the target object cluster. For example, preset conditions include application A, location A, within 3 days, and registered account. The target object cluster can be limited by application A, location A, and registered account, and the target object cluster must simultaneously meet the following conditions: an account of application A, the account login location is location A, and the account is a registered account. Correspondingly, the target data source is the operation data generated by the target object cluster within 3 days.

[0060] In one embodiment, before extracting preset conditions from the received profile query request and determining the target data source that meets the preset conditions in response to the received profile query request, the method may further include the following steps: receiving the profile query request, wherein the profile query request is a real-time request or a timed request generated based on operation and maintenance requirements information and / or the operation status of the server cluster involved in the processing link.

[0061] Profile query requests can be scheduled requests, with the server responding to these requests to retrieve target profile data for the target object cluster. Profile query requests can also be real-time requests, with the server responding to these requests to retrieve target profile data for the target object cluster. This section provides scenarios where operational requirements and the operational status of the server clusters involved in the processing chain are considered when generating real-time or scheduled requests, improving the adaptability of the generated profile query requests.

[0062] The processing chain is data source-oriented. Each task stage in the processing chain is responsible for processing the data source flowing to its own stage. Specifically, the server cluster responsible for executing the tasks in that stage processes the relevant data. Through the processing chain, profile data of the object cluster corresponding to the data source can be obtained. Operational maintenance (O&M) requirements describe the O&M management needs for the profile. O&M management for the profile can include monitoring the processing chain and identifying and resolving issues based on monitoring. O&M requirements can indicate periodic full-chain O&M requirements or task-stage O&M requirements, which aim to trigger the generation of scheduled profile query requests. Periodic O&M improves the orderliness and standardization of O&M management, thereby facilitating anomaly detection and resolution, and ensuring the quality of profile data obtained based on the processing chain. O&M requirements can also indicate full-chain O&M requirements or task-stage O&M requirements for a specified future time, which aims to trigger the generation of profile query requests at the specified time (i.e., scheduled), thus improving the convenience of obtaining relevant profile data for O&M management. Operation and maintenance requirement information can also indicate the full-link operation and maintenance requirements or task-stage operation and maintenance requirements at the current time. This operation and maintenance requirement information is intended to trigger the generation of real-time profile query requests, which can improve the convenience and timeliness of obtaining relevant profile data for operation and maintenance management.

[0063] The operational status of the server clusters involved in the processing chain. The processing chain comprises multiple task stages. For each task stage, there is a server cluster responsible for executing the tasks of that stage. The operational status of the server clusters involved in the processing chain can be obtained based on the operational status of multiple individual server clusters. The operational status of a server cluster can include whether the server cluster is partially or completely offline, whether it is partially or completely powered off, and the load status of the server cluster, such as hardware parameters like CPU utilization, memory read speed, and disk I / O performance. When the data source of the profile data to be acquired includes operational data generated by the object cluster in the current time period, the operational status of the server clusters involved in the processing chain can affect the timeliness of acquiring the profile data. This is especially true for time periods with a large volume of newly added operational data that can serve as a data source.

[0064] In practical applications, the portrait data quality monitoring system may include a first type of server cluster and multiple second type server clusters involved in the processing link. The first type of server cluster is used to execute the data processing scheme provided in the embodiments of this application. The operation of the server clusters involved in the processing link affects the operation of the portrait data quality monitoring system and whether it can respond normally to portrait query requests that are real-time requests. Accordingly, it can be determined whether to generate a portrait query request as a real-time request or a portrait query request as a timed request.

[0065] S202: Obtain target profile data of the target object cluster to perform anomaly monitoring on the processing link based on the target profile data. The target profile data is used to characterize target stage indicator data for each task stage on the processing link of the target data source. The target stage indicator data includes a first type of indicator data and a second type of indicator data. The first type of indicator data is used to describe the data flow of the task stage, and the second type of indicator data is used to describe the task execution of the task stage.

[0066] In this embodiment, the server acquires target profile data of the target object cluster to perform anomaly monitoring on the processing link based on the target profile data. Anomaly monitoring on the processing link helps maintain profile quality. The target profile data is used to characterize target stage indicator data for each task stage on the processing link for the target data source. The target stage indicator data includes: a first type of indicator data describing the data flow of the task stage, and a second type of indicator data describing the task execution status of the task stage. The target data source is the input of the first task stage of the processing link, and the data output of the previous task stage is the data input of the next task stage. Each task stage processes the input data of its own task stage to achieve task execution. The first type of indicator data is data-dimensional indicator data, which may include one of the following: the amount of input data in this task stage (e.g., the number of input data rows), the amount of output data in this task stage (e.g., the number of output data rows), the amount of valid input data in this task stage, and the amount of abnormal input data in this task stage. The second type of indicator data is task execution-dimensional indicator data, which may include at least one of the following: timeliness information of task execution, task execution effect, and the running status of the server cluster executing the task. If the task involves pre-processing the input data, the task execution effect can be determined by comparing the pre-processed result with the input data. The operational status of the server cluster executing the task focuses on reflecting the cluster's operation while the task is being executed; the start and end times of the statistics can indicate the start and end of the task execution. It should be noted that the target data source consists of the operational data of the target object cluster. Different types of operational data can be processed differently, specifically reflected in the different processing of different types of operational data in the input data at different task stages. This also affects the granularity of the first and second types of indicator data.

[0067] The target profile data obtained for the target object cluster originates from the processing of operational data provided by the target object cluster. The obtained target profile data describes the multi-dimensional characteristics of the target object cluster, and can also be composed of multi-dimensional features. Target stage indicator data from each task stage in the processing chain are included within the scope of multi-dimensional features. These target stage indicator data can cover the features captured by integrating operational data and also provide features strongly correlated with the processing stages of the target data source. This more global object cluster profile provides holistic data for personnel at all stages of profile operation and maintenance management, thereby improving the accuracy and efficiency of operation and maintenance management. Personnel at each stage of profile operation and maintenance management often only focus on the indicator data of their own stage, neglecting the data flowing into and out of their stage, or lacking the conditions to focus on these related data. This application's embodiments treat each stage of profile operation and maintenance management as a task phase in a processing chain, vertically linking the phase indicator data of each task phase. This overcomes the limitations of phase indicator data from a single task phase, providing support for locating abnormal task phases through upstream and downstream indicator data (e.g., comparing data outflow from a previous task phase with data inflow from a subsequent task phase), thus facilitating efficient operation and maintenance management. Through vertical linking, the phase indicator data of each task phase are no longer isolated; they can be provided as a whole to personnel at all stages of profile operation and maintenance management, unifying the profile data viewing perspective for personnel at different stages and improving the efficiency of locating abnormal task phases. Simultaneously, it provides traceable data for troubleshooting and saves time and manpower costs, as well as the waste of human and system resources, caused by a lack of a unified perspective.

[0068] In one embodiment, such as Figure 3 As shown, the processing chain includes multiple task stages, the preset conditions include a preset time period, and the acquisition of target profile data of the target object cluster includes:

[0069] S301: Obtain multiple candidate stage indicator data that conform to the preset time period. The multiple candidate stage indicator data correspond one-to-one with the multiple task stages. The candidate stage indicator data comes from the reported data of the server cluster responsible for executing the task stage. The candidate stage indicator data is stage indicator data with a preset time window as the time statistical caliber. The length of the preset time window is less than or equal to the length of the preset time period.

[0070] S302: Obtain the target profile data based on the multiple candidate stage indicator data.

[0071] Referring to the example in step S201 above, "the preset conditions include application A, location A, within 3 days, and registered account", the preset time period is "within 3 days" here.

[0072] For example, if the preset time period is from the day before yesterday to yesterday, and the preset time window is 1 day (24 hours of a calendar day), then the stage indicator data falling within the preset time period is obtained from the stage indicator data of each task stage. Due to the preset time window, there may be multiple sets of stage indicator data falling within the preset time period for each task stage.

[0073] Preset time periods and preset windows are used to limit the same attribute of operational data, such as the generation time, reporting time, and acquisition time of operational data. The server cluster responsible for executing tasks in each task phase can determine phase indicator data using the preset time window as the time statistical caliber and report it to the server. The processing chain includes multiple task phases. For each task phase, there is a server cluster responsible for executing the tasks in that phase. The server receives and saves the phase indicator data of the preset time window dimension reported by multiple server clusters. When responding to a profile query request, it determines the phase indicator data of multiple preset time window dimensions that meet the requirements from the stored data to obtain the target profile data. This mechanism of introducing the reporting of phase indicator data of the preset time window dimension provides timeliness support for obtaining object cluster profiles of the preset time period dimension, which helps improve the accuracy and real-time performance of obtaining the target profile data. This allows for timely acquisition of reference profile data during profile operation and maintenance management, enabling the identification and resolution of problems. In practical applications, the preset time window can be flexibly set according to actual needs, such as 1 second, 1 hour, or 1 week.

[0074] Furthermore, obtaining multiple candidate stage indicator data that conform to the preset time period may include the following steps: for each task stage, obtaining stage indicator data that conforms to the preset time period from the storage module indicating the task stage, so as to obtain the multiple candidate stage indicator data.

[0075] For each task stage, stage indicator data conforming to a preset time period is retrieved from the storage module indicating the task stage. This stage indicator data, conforming to the preset time period, constitutes the target stage indicator data for the task stage. Combined with the aforementioned server cluster, stage indicator data is reported within a preset time window dimension. The server-side saves the reported data, allowing for task stage-level storage and maintenance. This helps reduce the difficulty of locating stage indicator data conforming to the preset time period for a specific task stage, thereby improving the efficiency of obtaining target profile data. Storing and maintaining stage indicator data separately for different task stages facilitates further analysis of the stage indicator data for a particular task stage.

[0076] In one embodiment, the processing chain consists of a data source acquisition stage, a data cleaning stage, a data preprocessing stage, and a data statistics stage, following the flow direction of the target data source. Acquiring the target profile data of the target object cluster may include the following steps: First, acquiring first-stage indicator data indicating the data source acquisition stage, whereby the first-stage indicator data indicates the original quantity of the acquired target data source, data quality, and the timeliness information of executing the data source acquisition task; then, acquiring second-stage indicator data for the data cleaning stage, whereby the second-stage indicator data indicates the data inflow, data outflow, and data cleaning status of the data cleaning stage. The data preprocessing stage includes the following steps: obtaining the data inflow, data outflow, and timeliness information for performing data cleaning tasks; obtaining the third-stage indicator data for the data preprocessing stage, which indicates the data inflow, data outflow, and timeliness information for performing data preprocessing tasks; obtaining the fourth-stage indicator data for the data statistics stage, which indicates the data inflow, data statistics results, and the operational status of the server cluster performing data statistics tasks; and finally, obtaining the target profile data based on the first-stage indicator data, the second-stage indicator data, the third-stage indicator data, and the fourth-stage indicator data.

[0077] This section introduces the data source acquisition stage, data cleaning stage, data preprocessing stage, and data statistics stage used to construct the processing chain, and provides specific data dimension indicators and task execution dimension indicators for each task stage. The processing chain provided here can enrich the feature granularity of the object cluster profile, thereby improving the profile quality.

[0078] The data source acquisition stage acquires data from the data source and inputs processed data into the data cleaning stage. This processing primarily focuses on data with obvious anomalies in the data source. The data cleaning stage cleans the data input from the data source acquisition stage and inputs the cleaned data into the data preprocessing stage. This cleaning can be achieved using preset cleaning rules, which focus on filtering out dirty data that might affect the generation of the object cluster profile. The data preprocessing stage preprocesses the data input from the data cleaning stage and inputs the preprocessed data into the data statistics stage. This preprocessing aims to integrate all or part of the input data to obtain intermediate data. Intermediate data can include labels, merged information, etc., for all or part of the input data. Preprocessing can be implemented through simple calculations or using AI models. The data statistics stage stores and statistically analyzes the data input from the data preprocessing stage. Statistical methods include, but are not limited to, cumulative statistics, mean statistics, median statistics, variance statistics, etc. If the aforementioned preprocessing applies to all input data, then the input data to be statistically analyzed is intermediate data. If the aforementioned preprocessing applies to part of the input data, then the input data to be statistically analyzed includes intermediate data and the aforementioned unprocessed portion of the input data. For example, the data source includes the action logs of users 1-N, i.e., action log set 1. After data cleaning, action log set 2 is obtained. A portion of the action logs in action log set 2 is preprocessed to obtain corresponding intermediate data. The input data for the data statistics stage can include this intermediate data, as well as the unprocessed portion of the action logs in action log set 2. Taking a user's use of a video application as an example, a user's action log is the action data generated by the user's use of the video application within a time period, which can include actions such as viewing viewing history, watching a video, and downloading a video. Here, the action logs of N users together constitute action log set 1.

[0079] For example, a portrait data quality monitoring system may include a first type of server cluster and multiple second type of server clusters involved in the processing link. The first type of server cluster is used to execute the data processing scheme provided in the embodiments of this application. The portrait data quality monitoring system can provide object cluster portrait services for application A, and the multiple second type of server clusters can access the backend server of application A to obtain the user's operation log of application A. (Reference) Figure 5The aforementioned raw quantities correspond to the number of record rows in the diagram, the aforementioned data quality corresponds to the number of valid rows in the diagram, and the data delivery time in the diagram represents the time between the generation of the data and its arrival at the backend server. The aforementioned data cleaning status corresponds to the number of abnormal rows in the diagram. In practical applications, the data statistics stage can be divided into a data import sub-stage, a data storage sub-stage, and a data statistics sub-stage. For each sub-stage, there is a server cluster responsible for executing the tasks of that sub-stage. The operational status of the server cluster executing the data import task can include the import time in the diagram. The operational status of the server cluster executing the data storage task can include the query time and concurrency in the diagram. The operational status of the server cluster executing the data statistics task can include the statistical timeliness and concurrency in the diagram. Figure 5 As shown, this embodiment of the application links the indicators of each task stage to form a global monitoring of the profile dimension. Here, it is no longer a single indicator for each task stage, but a summary indicator, and the perspective of the indicator is entirely focused on the profile features.

[0080] In one embodiment, such as Figure 4 As shown, after obtaining the target profile data of the target object cluster, the method further includes:

[0081] S401: Embed the target image data into a preset template to obtain page data;

[0082] S402: Send the page data to a preset terminal so that the preset terminal displays the target page based on the page data.

[0083] By displaying the target page through pre-set terminals, staff can gain a more intuitive understanding of the target profile data. The use of pre-set templates improves the standardization and readability of the front-end display page; this can be used as a reference. Figure 6 .

[0084] In one embodiment, the processing link includes multiple task stages. After acquiring the target profile data of the target object cluster, the method may further include the following steps: if the target stage indicator data of any task stage in the multiple task stages does not meet preset requirements, determine an abnormal task stage from the multiple task stages based on the target profile data; or, if the target stage indicator data of any task stage in the multiple task stages does not meet preset requirements, determine a reference condition that is related to the preset condition, determine a reference data source that meets the reference condition, acquire reference profile data of the reference object cluster, and determine an abnormal task stage from the multiple task stages based on the target profile data and the reference profile data, wherein the reference data source is composed of the operation data of the reference object cluster.

[0085] This provides a first method for determining the stage of an abnormal task using the target profile data itself, and a second method for determining the stage of an abnormal task using both the target profile data and the reference profile data. This provides a feasible approach to determining the stage of an abnormal task. The target profile data provides traceable data for troubleshooting, which can help to promptly identify and resolve problems in the operation and maintenance management of the profile, thereby ensuring the quality of the object cluster profile.

[0086] If the processing chain is divided into three stages according to the flow of the target data source, the target profile data includes target stage indicator data for the first stage, the second stage, and the third stage. If the target stage indicator data for the third stage does not meet the preset requirements, it may be due to a problem with the processing of the input data in the third stage, or it may be due to a problem with the input data itself. For the former, the problem can be mainly found in the task execution of the third stage itself; for the latter, the problem can be mainly found in the task execution of the second stage or even the first stage, which are related to the provision of the input data.

[0087] Therefore, taking the first approach as an example, at least one abnormal task stage can be identified from the first, second, and third stages based on the target stage indicator data of the first stage, the target stage indicator data of the second stage, and the target stage indicator data of the third stage.

[0088] Taking the second method as an example, reference conditions that are related to preset conditions can be identified, reference data sources that meet the reference conditions can be identified, and reference profile data of the reference object cluster can be obtained. The reference data source consists of the operation data of the reference object cluster, and the reference profile data includes reference stage indicator data for the first stage, reference stage indicator data for the second stage, and reference stage indicator data for the third stage. Then, based on the target stage indicator data for the first stage, the target stage indicator data for the second stage, the target stage indicator data for the third stage, the reference stage indicator data for the first stage, the reference stage indicator data for the second stage, and the reference stage indicator data for the third stage, at least one abnormal task stage can be identified from the first, second, and third stages. Combining the example of "preset conditions include application A, location A, within 3 days, and registered account" in the aforementioned step S201, the reference conditions can be obtained by associating and replacing the condition items in the preset conditions. For example, if there are inclusion and containment relationships within 1 day and 3 days, the reference data source can be determined by specifying "reference conditions include application A, location B, within 1 day, and registered account". The reference profile data of the reference object cluster can provide a reference for identifying the abnormal task stage. Similarly, if the traffic of application A in location B and location A is similar, the reference data source can be determined by specifying "reference conditions include application A, location B, within 3 days, and registered account". The reference profile data of the reference object cluster can also provide a reference for identifying the abnormal task stage. The application of reference profile data can be reflected in comparing the data output of the previous task stage with the data input of the next task stage, and comparing whether the target stage indicator data and the reference stage indicator data of the same task stage meet the requirements. The application of reference profile data can further improve the accuracy of identifying the abnormal task stage. It should be noted that the process of "determining the reference data source that meets the reference conditions and obtaining the reference profile data of the reference object cluster" can refer to the records related to "determining the target data source that meets the preset conditions" and "obtaining the target profile data of the target object cluster", and will not be elaborated further.

[0089] After identifying the abnormal task phase using the first and second methods described above, an alarm notification can be generated based on the abnormal task phase. This alarm notification instructs the relevant personnel to perform operational fault handling on the abnormal task phase. Then, the alarm notification is sent to the client. The client to which the alarm notification is sent is staff members. The alarm notification promptly guides these staff members to handle the operational faults of the abnormal task phase. Operational fault handling includes, but is not limited to, checking the algorithm logic used in the abnormal task phase and repairing the task execution of the abnormal task phase. Of course, the repair targets also include the identified target phase indicator data with anomalies, and the object cluster profile related to the target phase indicator data.

[0090] As can be seen from the technical solutions provided in the embodiments of this application above, the embodiments of this application provide a more global object cluster profile. This profile is used to characterize the target stage indicator data of each task stage in the processing link of the target data source. The target stage indicator data includes both data-dimensional indicator data and task execution-dimensional indicator data. Through such a profile, anomaly monitoring of the processing link can be performed, thereby achieving accurate and effective profile quality monitoring, and thus providing more targeted services to relevant users. Such a profile can provide holistic data for profile operation and maintenance management, improving the accuracy and efficiency of operation and maintenance management.

[0091] This application embodiment also provides a data processing method applied to a data processing system. The data processing system includes a first type of server cluster and multiple second type of server clusters. The first type of server cluster is used to execute the data processing method described in steps S201-S202 above. The multiple second type of server clusters are respectively responsible for executing the tasks of each task stage in the processing link. The method includes: for each task stage, the second type of server cluster corresponding to the task stage generates stage indicator data to be reported with a preset time window as the time statistical caliber, and reports the stage indicator data to be reported to the first type of server cluster.

[0092] By introducing a reporting mechanism for stage indicator data with a preset time window dimension, timely support is provided for obtaining object cluster profiles with a preset time period (the length of the preset time window is less than or equal to the length of the preset time period), which helps to improve the accuracy and real-time performance of obtaining target profile data.

[0093] In one embodiment, where the processing chain includes a data source acquisition stage, a data preprocessing stage, and a data statistics stage, the second type of server cluster corresponding to the data source acquisition stage uses operational data storage to store the stage indicator data of the data source acquisition stage; the second type of server cluster corresponding to the data preprocessing stage uses a data detail layer to store the stage indicator data of the data preprocessing stage; and the second type of server cluster corresponding to the data statistics stage uses an online storage module to store the stage indicator data of the data statistics stage. The first type of server cluster uses a preset relational database management system to store the reported data of the multiple second type of server clusters.

[0094] For example, the data processing system of the portrait data quality monitoring system may include a first type of server cluster and multiple second type of server clusters involved in the processing link. The first type of server cluster is used to execute the data processing methods described in steps S201-S202 above. The multiple second type of server clusters may report stage indicator data to the first type of server cluster. The first type of server cluster may also request stage indicator data from the multiple second type of server clusters. The request method may include scheduled capture, manual synchronization, and real-time capture. Scheduled capture tasks can be implemented using Spark or Oozie, and scheduled capture tasks can be triggered using the crontab command.

[0095] For the processing chain, relevant data can be stored using a data warehouse, such as MySQL. The data source acquisition stage can utilize ODS to store relevant data, specifically TDW and ATTA. The data preprocessing stage can utilize DWD to store relevant data. The data storage sub-stage can utilize Redis, BDB, YottaDB, or KeeWiDB to store relevant data. The data storage sub-stage can also store relevant data online. The data stored in the ODS includes not only the stage indicator data from the data source acquisition stage but also primarily metadata. Metadata can come from the operational data of the target object cluster and also from the identification information of the data processing system (such as business scenarios, task information, resource information, etc., as described later). The identification information of the data processing system helps to understand the basic situation of the data processing system. Metadata includes entity information and the relationship information between entities. Entity information can include business scenarios, task information, resource information, feature information, ID information, etc. Business scenarios can indicate the application of the profile data quality monitoring system service or the purpose of the application, such as the profile data quality monitoring system providing object cluster profile services for application A for recommendation, growth, commercialization, or advertising businesses. Task information describes the tasks at each stage of the processing chain, such as whether the task is offline or real-time, its content, and the object responsible for it. Resource information can include computing resource information, offline resource information, and resource ownership information. Feature information determines the granularity of the profile and can include stage indicator data used to build the object cluster profile. ID information can include ID type and ID mapping relationship. A user's login account can be one type of ID, and a user's device ID can be another type. One login account may correspond to multiple device IDs, which is an ID mapping relationship. Relationship information between entities can be reflected in the association between tasks and resources (e.g., task i uses computing resource j), the association between tasks and features (e.g., task i generates feature k), and the association between business and features (e.g., business l uses feature k). In addition to stage indicator data, the manually synchronized objects mentioned above can also include some non-stage indicator data in the metadata, especially some data that does not change very frequently. For example, what data cleaning tasks are currently available, the input data of the data cleaning stage, the data storage media used, and the storage space size, etc., which are implemented through configurable scripts.

[0096] The first type of server cluster can be maintained using MySQL for reporting data and / or requesting data. Below is an example of obtaining an object cluster profile. If the object cluster profile consists of stage indicator data from the aforementioned data source acquisition stage, data cleaning stage, data preprocessing stage, and data statistics stage, the object cluster profile indicates a specific business, location, and ID type. The object cluster characteristic can be set as the key, and the object cluster profile as the value:

[0097] value = profile feature join metadata on feature id

[0098] join business ON feature id

[0099] Join multiple locations to deploy on profile features

[0100] join id type

[0101] join import subtask on import subtask id

[0102] join data preprocessing task on data preprocessing task id

[0103] join data cleaning task on data cleaning task id

[0104] Joining the data source retrieves the task ID from the ON data source.

[0105] The first type of server cluster introduces tables when performing data maintenance, and the object cluster profile can be obtained through multi-table join queries.

[0106] The first type of server cluster can periodically generate profile query requests to obtain and store profile data at fixed intervals. It also provides profile data visualization. The backend server architecture for profile data visualization can be built using Spring Boot (an open-source application framework) and connect to the aforementioned MySQL database via JDBC (an application programming interface) to obtain the underlying data. A backend dashboard display can be found here. Figure 6 , 7 The dashboard supports retrieving full-chain periodic metrics for each feature based on business, location, ID type, feature name, and time dimension.

[0107] This application also provides a data processing apparatus, such as... Figure 8 As shown, the data processing device 80 includes:

[0108] Response module 801: In response to a received profile query request, extract preset conditions from the profile query request and determine a target data source that meets the preset conditions, wherein the target data source consists of operation data of the target object cluster;

[0109] Acquisition module 802: used to acquire target profile data of the target object cluster, so as to perform anomaly monitoring on the processing link based on the target profile data. The target profile data is used to characterize the target stage indicator data of each task stage on the processing link for the target data source. The target stage indicator data includes a first type of indicator data and a second type of indicator data. The first type of indicator data is used to describe the data flow of the task stage, and the second type of indicator data is used to describe the task execution of the task stage.

[0110] It should be noted that the apparatus and method embodiments described in the device embodiments are based on the same inventive concept.

[0111] In some embodiments, the functions or modules of the apparatus provided in this application can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0112] This application also provides a computer-readable storage medium storing at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the above-described method. The computer-readable storage medium may be a non-volatile computer-readable storage medium.

[0113] This application also provides an electronic device, which includes at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the at least one processor to implement the above method.

[0114] Electronic devices can be provided as terminals, servers, or other forms of devices.

[0115] Figure 9 A block diagram of an electronic device according to an embodiment of this application is shown. For example, electronic device 1900 may be provided as a server. (Refer to...) Figure 9The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.

[0116] Electronic device 1900 may also include a power supply component 1926 configured to perform power management of electronic device 1900, a wired or wireless network interface 1950 configured to connect electronic device 1900 to a network, and an input / output (I / O) interface 1958. Electronic device 1900 can operate on an operating system stored in memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.

[0117] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of an electronic device 1900 to perform the above-described method.

[0118] This application may be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this application.

[0119] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0120] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0121] The computer program instructions used to perform the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C+, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing the status information of the computer-readable program instructions. These electronic circuits can execute the computer-readable program instructions to implement various aspects of this application.

[0122] Various aspects of this application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0123] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0124] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which includes one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions specified in the blocks may occur in a different order than those specified in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0126] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technological improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A data processing method, characterized in that, The method includes: In response to a received profile query request, preset conditions are extracted from the profile query request, and a target data source that meets the preset conditions is determined. The target data source consists of operation data of the target object cluster. Obtain target profile data of the target object cluster to perform anomaly monitoring on the processing link based on the target profile data. The target profile data is used to characterize target stage indicator data for each task stage on the processing link of the target data source. The target stage indicator data includes a first type of indicator data and a second type of indicator data. The first type of indicator data is used to describe the data flow of the task stage, and the second type of indicator data is used to describe the task execution of the task stage.

2. The method according to claim 1, characterized in that, The processing chain includes multiple task stages, the preset conditions include a preset time period, and the acquisition of target profile data of the target object cluster includes: Obtain multiple candidate stage indicator data that conform to the preset time period. The multiple candidate stage indicator data correspond one-to-one with the multiple task stages. The candidate stage indicator data comes from the reported data of the server cluster responsible for executing the task stage. The candidate stage indicator data is stage indicator data with a preset time window as the time statistical caliber. The length of the preset time window is less than or equal to the length of the preset time period. The target profile data is obtained based on the multiple candidate stage indicator data.

3. The method according to claim 2, characterized in that, The step of obtaining multiple candidate stage indicator data that conform to the preset time period includes: For each task stage, stage indicator data that conforms to the preset time period is obtained from the storage module indicating the task stage to obtain the multiple candidate stage indicator data.

4. The method according to claim 1, characterized in that, In the processing chain, according to the flow direction of the target data source, the stages are: data source acquisition, data cleaning, data preprocessing, and data statistics. Acquiring the target profile data of the target object cluster includes: Obtain first-stage indicator data indicating the data source acquisition phase, wherein the first-stage indicator data indicates the original quantity, data quality, and timeliness information of the target data source acquired and the execution of the data source acquisition task; Obtain the second-stage indicator data of the data cleaning stage. The second-stage indicator data indicates the data inflow, data outflow, data cleaning status, and timeliness information of the data cleaning task in the data cleaning stage. Obtain the third-stage indicator data of the data preprocessing stage, which indicates the data inflow, data outflow, and timeliness information of the data preprocessing task in the data preprocessing stage. Obtain the fourth-stage indicator data of the data statistics phase, which indicates the data inflow, data statistics results, and the operating status of the server cluster executing the data statistics task in the data statistics phase. The target profile data is obtained based on the first-stage indicator data, the second-stage indicator data, the third-stage indicator data, and the fourth-stage indicator data.

5. The method according to any one of claims 1-4, characterized in that, After obtaining the target profile data of the target object cluster, the method further includes: The target profile data is embedded into a preset template to obtain page data; The page data is sent to a preset terminal so that the preset terminal displays the target page based on the page data.

6. The method according to claim 1, characterized in that, The processing chain includes multiple task stages. After acquiring the target profile data of the target object cluster, the method further includes: If the target stage indicator data of any of the multiple task stages does not meet the preset requirements, an abnormal task stage is determined from the multiple task stages based on the target profile data; or... If the target stage indicator data of any of the multiple task stages does not meet the preset requirements, a reference condition that is related to the preset condition is determined, a reference data source that meets the reference condition is determined, reference profile data of the reference object cluster is obtained, and an abnormal task stage is determined from the multiple task stages based on the target profile data and the reference profile data. The reference data source is composed of the operation data of the reference object cluster.

7. The method according to claim 6, characterized in that, The method further includes: An alarm notification is generated based on the abnormal task phase, and the alarm notification is used to indicate that the abnormal task phase should be handled for maintenance faults. Send the alarm notification to the client.

8. The method according to claim 1, characterized in that, Before extracting preset conditions from the received profile query request and determining the target data source that meets the preset conditions in response to the received profile query request, the method further includes: The profile query request is received. The profile query request is a real-time request or a timed request generated based on operation and maintenance requirements information and / or the operation status of the server cluster involved in the processing link.

9. A data processing method, characterized in that, The data processing system is applied to a data processing system comprising a first type of server cluster and multiple second type server clusters. The first type of server cluster is used to execute the data processing method as described in any one of claims 1-8, and the multiple second type server clusters are respectively responsible for executing tasks at each task stage of the processing link. The method includes: For each task stage, the second type of server cluster corresponding to the task stage generates stage indicator data to be reported using a preset time window as the time statistical caliber, and reports the stage indicator data to be reported to the first type of server cluster.

10. The method according to claim 9, characterized in that, In the case that the processing chain includes a data source acquisition stage, a data preprocessing stage, and a data statistics stage, The second type of server cluster corresponding to the data source acquisition phase uses operational data storage to store the phase indicator data of the data source acquisition phase; The second type of server cluster corresponding to the data preprocessing stage uses the data detail layer to store the stage indicator data of the data preprocessing stage; The second type of server cluster corresponding to the data statistics phase uses an online storage module to save the phase indicator data of the data statistics phase.

11. The method according to claim 9, characterized in that, The first type of server cluster uses a preset relational database management system to store the reported data of the multiple second type of server clusters.

12. A data processing apparatus, characterized in that, The device includes: Response module: Used to respond to a received profile query request, extract preset conditions from the profile query request, and determine the target data source that meets the preset conditions. The target data source consists of operation data of the target object cluster. Acquisition Module: Used to acquire target profile data of the target object cluster, so as to perform anomaly monitoring on the processing link based on the target profile data. The target profile data is used to characterize the target stage indicator data of each task stage on the processing link for the target data source. The target stage indicator data includes a first type of indicator data and a second type of indicator data. The first type of indicator data is used to describe the data flow of the task stage, and the second type of indicator data is used to describe the task execution of the task stage.

13. An electronic device, characterized in that, The electronic device includes at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores at least one instruction or at least one program, the at least one instruction or at least one program being loaded and executed by the at least one processor to implement the data processing method as described in any one of claims 1-8, or the data processing method as described in any one of claims 9-11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction or at least one program, which is loaded and executed by a processor to implement the data processing method as described in any one of claims 1-8 or the data processing method as described in any one of claims 9-11.

Citation Information

Patent Citations

  • Portrait analysis method and device based on big data, computer device and storage medium

    CN110363387A

  • Method, device and equipment for integrating customer portrait indexes

    CN112579655A