Metadata acquisition method and device

By obtaining historical execution information and current time in the metadata acquisition system, determining the data acquisition queue, and planning the acquisition task based on pipeline resource limitations, the problems of high resource consumption and low acquisition efficiency in traditional methods are solved, and efficient and stable metadata acquisition is achieved.

CN120066692APending Publication Date: 2025-05-30HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311609167.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional metadata acquisition methods ignore resource consumption, resulting in high CPU or memory load on the acquisition system and source application system, affecting system stability, and ineffective and inefficient acquisition, resulting in low acquisition efficiency.

Method used

By obtaining historical execution information, current time and task type priority, determining the data acquisition queue, accurately collect based on data characteristics, and limiting the number of acquisition tasks through pipeline resources to avoid system suspension and blockage and reduce resource consumption.

Benefits of technology

Accurate acquisition based on data features is realized, avoiding system suspension and blockage, reducing resource consumption for real-time data acquisition, and improving acquisition efficiency and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066692A_ABST
    Figure CN120066692A_ABST
Patent Text Reader

Abstract

The invention discloses a metadata collection method and device.The method comprises the steps that historical execution information, current time and task type priorities are obtained, and the historical execution information comprises first execution information of M first collection tasks, second execution information of X second collection tasks and third execution information of P third collection tasks, m is an integer greater than 1, X is an integer greater than 1, and P is an integer greater than 1; according to the first execution information, determining time statistical information and predicted consumption duration of the M first acquisition tasks; determining a data acquisition queue according to the second execution information, the third execution information, the current time, the task type priority, the time statistical information and the predicted consumed duration; and performing data acquisition based on the data acquisition queue. By adopting the method and the device, accurate acquisition can be performed based on data features, resource consumption of real-time data acquisition is reduced, and acquisition efficiency and system stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for metadata collection. Background Art

[0002] With the rapid development of computer technology and the continuous deepening of enterprise digitization, enterprises' demand for metadata collection is constantly increasing. Traditional metadata collection mainly includes the timed collection method and the real-time collection method. However, these traditional methods ignore the actual environmental resource consumption in the live network, and call the collector to collect data through a fixed task scheduling plan, resulting in high loads on the central processing unit (CPU) or memory of the collection system and the source application system, and even affecting the source jobs and causing system anomalies; moreover, these traditional methods do not consider data characteristics, and there may be situations of invalid and inefficient collection, resulting in low collection efficiency. Summary of the Invention

[0003] Embodiments of the present application provide a method and device for metadata collection, which can perform precise collection based on data characteristics, plan the number of collection tasks through resource restrictions, avoid system suspension and blockage caused by a large number, reduce the resource consumption of real-time data collection, and improve the collection efficiency and system stability.

[0004] In a first aspect, embodiments of the present application provide a method for metadata collection, including:

[0005] Obtain historical execution information, current time, and task type priorities. The historical execution information includes the first execution information of M first collection tasks, the second execution information of X second collection tasks, and the third execution information of P third collection tasks. The first collection tasks are completed tasks, the second collection tasks are running tasks, and the third collection tasks are unrun tasks. M is an integer greater than 1, X is an integer greater than 1, and P is an integer greater than 1.

[0006] Determine the time statistical information and estimated consumption duration of the M first collection tasks according to the first execution information.

[0007] Determine a first queue of tasks to be executed according to the second execution information, current time, task type priorities, time statistical information, and estimated consumption duration.

[0008] Determine a second queue of tasks to be executed according to the third execution information, current time, task type priorities, time statistical information, and estimated consumption duration.

[0009] Determine a data collection queue according to the first queue of tasks to be executed, the second queue of tasks to be executed, the second execution information, and the third execution information.

[0010] Data collection is performed based on the data collection queue.

[0011] Queue planning is carried out by combining historical execution information, current time, and task type priority to obtain the first to-be-executed queue and the second to-be-executed queue. Then, the pipeline resource limit is determined through the second execution information and the third execution information. Next, the first to-be-executed queue and the second to-be-executed queue are planned based on the pipeline resource limit, and thus the data collection queue is determined. Through the data collection queue, accurate collection can be performed based on data characteristics, and the number of collection tasks is planned through the pipeline resource limit, which can avoid system suspension and blockage caused by a large number, reduce the resource consumption of real-time data collection, and improve the collection efficiency and system stability.

[0012] In a possible design, the first execution information includes the actual start time and actual completion time of each of the M first collection tasks; according to the actual start time and actual completion time, time statistical information is calculated, and the time statistical information includes the expected value and standard deviation; the expected value is determined as the expected consumption duration. By determining the time statistical information and the expected consumption duration, it is convenient for subsequent calculation of the planned start time of the task, planning of the first to-be-executed queue and the second to-be-executed queue, which is beneficial to reducing the resource consumption of real-time data collection.

[0013] In another possible design, the second execution information includes the task type, task source, and first preset start time of each of the X second collection tasks; according to the task type and task source of each of the X second collection tasks, the first pipeline information is determined, and the first pipeline information includes the collection pipeline type and pipeline parallelism of each of the X second collection tasks; according to the first preset start time, the first pipeline information, the current time, the task type priority, the time statistical information, and the expected consumption duration, the first to-be-executed queue is determined. By determining the collection pipeline type and pipeline parallelism of each second collection task, it is convenient for subsequent planning of the first to-be-executed queue, which is beneficial to achieving accurate collection based on data characteristics and avoiding ineffective and inefficient collection.

[0014] In another possible design, according to the first preset start time and standard deviation of each of the X second collection tasks, the first planned start time of each of the X second collection tasks is determined; according to the first preset start time, the first pipeline information, the first planned start time, the current time, the task type priority, and the expected consumption duration, the first to-be-executed queue is determined. By determining the first planned start time of each second collection task, it is convenient for subsequent screening of the X second collection tasks, which is beneficial to planning the number of collection tasks through resource limits and avoiding system suspension and blockage caused by a large number.

[0015] In another possible design, the first planned start time satisfies:

[0016] T 1,i = H 1,i - k 1 × S;

[0017] where T 1,i represents the first planned start time of the i-th second acquisition task among X second acquisition tasks, H 1,i represents the first preset start time of the i-th second acquisition task among X second acquisition tasks, S represents the standard deviation, k 1 represents the empirical parameter, and i is an integer greater than 0 and less than or equal to X.

[0018] In another possible design, X second acquisition tasks are screened according to the first preset start time, the first pipeline information, and the current time to obtain Y second acquisition tasks, where Y is an integer greater than 1 and less than or equal to X; according to the Y second acquisition tasks, the first pipeline information, the first planned start time, the task type priority, and the estimated consumption duration, a first to-be-executed queue is determined. Screening to obtain Y second acquisition tasks that meet the conditions is beneficial to planning the number of acquisition tasks through resource limitations and avoiding system suspension and congestion when the quantity is large.

[0019] In another possible design, the current time is subtracted from the first preset start time of each second acquisition task among the X second acquisition tasks to obtain the waiting duration of each second acquisition task among the X second acquisition tasks; if the first waiting duration among the X waiting durations is less than or equal to the first preset threshold, the second acquisition task corresponding to the first waiting duration is used as one of the Y second acquisition tasks. Screening to obtain Y second acquisition tasks with a waiting duration less than or equal to the first preset threshold is beneficial to planning the number of acquisition tasks through resource limitations and avoiding system suspension and congestion when the quantity is large.

[0020] In another possible design, the Y second acquisition tasks are divided according to the acquisition pipeline type to obtain at least one to-be-sorted task subset; according to the first planned start time, the task type priority, and the estimated consumption duration, at least one to-be-sorted task subset is sorted to obtain a first to-be-executed queue corresponding to each acquisition pipeline type. By determining the first to-be-executed queue corresponding to each acquisition pipeline type, accurate acquisition based on data characteristics can be achieved, and ineffective and inefficient acquisition can be avoided.

[0021] In another possible design, the third execution information includes the task type, task source, and second preset start time of each of the P third collection tasks; according to the task type and task source of each of the P third collection tasks, the second pipeline information is determined, and the second pipeline information includes the collection pipeline type and pipeline parallelism of each of the P third collection tasks; according to the second preset start time, second pipeline information, current time, task type priority, time statistical information, and estimated consumption duration, the second to-be-executed queue is determined. By determining the collection pipeline type and pipeline parallelism of each third collection task, subsequent planning of the second to-be-executed queue is facilitated, which is conducive to achieving precise collection based on data characteristics and avoiding ineffective and inefficient collection.

[0022] In another possible design, according to the second preset start time and standard deviation of each of the P third collection tasks, the second planned start time of each of the P third collection tasks is determined; the estimated consumption duration is added to the second planned start time of each of the P third collection tasks to obtain the planned completion time of each of the P third collection tasks; according to the second preset start time, second pipeline information, planned completion time, current time, task type priority, and estimated consumption duration, the second to-be-executed queue is determined. By determining the second planned start time of each third collection task, subsequent screening of the P third collection tasks is facilitated, which is conducive to planning the number of collection tasks through resource limitations and avoiding system suspension and blockage when the quantity is large.

[0023] In another possible design, the second planned start time satisfies:

[0024] T 2,j =H 2,j -k 1 ×S;

[0025] wherein, T 2,j represents the second planned start time of the j-th third collection task among the P third collection tasks, H 2,j represents the second preset start time of the j-th third collection task among the P third collection tasks, S represents the standard deviation, k 1 represents the empirical parameter, and j is an integer greater than 0 and less than or equal to P.

[0026] In another possible design, based on the second preset start time, the second pipeline information, and the current time, P third acquisition tasks are screened to obtain Q third acquisition tasks, where Q is an integer greater than 1 and less than or equal to P; according to the Q third acquisition tasks, the second pipeline information, the planned completion time, the task type priority, and the estimated consumption duration, a second to-be-executed queue is determined. Screening to obtain Q third acquisition tasks that meet the conditions is beneficial for planning the number of acquisition tasks through resource limitations, and avoiding system suspension and congestion when the quantity is large.

[0027] In another possible design, the current time is subtracted from the second preset start time of each third acquisition task among the P third acquisition tasks to obtain the waiting duration of each third acquisition task among the P third acquisition tasks; if the second waiting duration among the P waiting durations is less than or equal to the second preset threshold, the third acquisition task corresponding to the second waiting duration is used as one of the Q third acquisition tasks. Screening to obtain Q third acquisition tasks with a waiting duration less than or equal to the second preset threshold is beneficial for planning the number of acquisition tasks through resource limitations, and avoiding system suspension and congestion when the quantity is large.

[0028] In another possible design, according to the acquisition pipeline type, the Q third acquisition tasks are divided to obtain at least one adjacent task subset; according to the planned completion time, the task type priority, and the estimated consumption duration, the at least one adjacent task subset is sorted to obtain the second to-be-executed queue corresponding to each acquisition pipeline type. By determining the second to-be-executed queue corresponding to each acquisition pipeline type, accurate acquisition based on data characteristics can be achieved, and ineffective and inefficient acquisition can be avoided.

[0029] In another possible design, according to the first pipeline information and the second pipeline information, the maximum pipeline parallelism of each acquisition pipeline type is determined; for the first to-be-executed queue corresponding to each acquisition pipeline type, the maximum pipeline parallelism is subtracted from the pipeline parallelism of the m-th second acquisition task in the first to-be-executed queue to obtain the remaining pipeline resources corresponding to each acquisition pipeline type, where m is an integer greater than 0 and less than or equal to Y; according to the remaining pipeline resources, the first to-be-executed queue, the second to-be-executed queue, and the second pipeline information, a data acquisition queue is determined. By determining the data acquisition queue, accurate acquisition based on data characteristics can be achieved, and by planning the number of acquisition tasks through pipeline resource limitations, system suspension and congestion when the quantity is large can be avoided, the resource consumption of real-time data acquisition can be reduced, and the acquisition efficiency and system stability can be improved.

[0030] In another possible design, for the second to-be-executed queue corresponding to each type of acquisition pipeline, it is determined whether the pipeline parallelism of the third acquisition task in the second to-be-executed queue is less than or equal to the remaining pipeline resources; if the first pipeline parallelism of the nth third acquisition task in the second to-be-executed queue is less than or equal to the remaining pipeline resources corresponding to the first pipeline parallelism, then the mth second acquisition task and the nth third acquisition task in the first to-be-executed queue are successively arranged into the data acquisition queue, where n is an integer greater than 0 and less than or equal to Q. By planning the number of acquisition tasks through pipeline resource constraints, it is possible to avoid system suspension and blockage when the quantity is large, reduce the resource consumption of real-time data acquisition, and improve the acquisition efficiency and stability.

[0031] In a second aspect, an embodiment of the present application provides a metadata acquisition device, including:

[0032] An acquisition module, configured to acquire historical execution information, current time, and task type priorities. The historical execution information includes the first execution information of M first acquisition tasks, the second execution information of X second acquisition tasks, and the third execution information of P third acquisition tasks. The first acquisition tasks are completed tasks, the second acquisition tasks are running tasks, and the third acquisition tasks are unrun tasks. M is an integer greater than 1, X is an integer greater than 1, and P is an integer greater than 1.

[0033] A processing module, configured to determine the time statistical information and the expected consumption duration of the M first acquisition tasks according to the first execution information.

[0034] The processing module is further configured to determine the first to-be-executed queue according to the second execution information, current time, task type priorities, time statistical information, and expected consumption duration.

[0035] The processing module is further configured to determine the second to-be-executed queue according to the third execution information, current time, task type priorities, time statistical information, and expected consumption duration.

[0036] The processing module is further configured to determine the data acquisition queue according to the first to-be-executed queue, the second to-be-executed queue, the second execution information, and the third execution information.

[0037] The processing module is further configured to perform data acquisition based on the data acquisition queue.

[0038] In a possible design, the first execution information includes the actual start time and actual completion time of each of the M first acquisition tasks.

[0039] In another possible design, the processing module is further configured to calculate the time statistical information according to the actual start time and actual completion time. The time statistical information includes an expected value and a standard deviation; and determine the expected consumption duration as the expected value.

[0040] In another possible design, the second execution information includes the task type, task source, and first preset start time of each of the X second acquisition tasks.

[0041] In another possible design, the processing module is further configured to determine first pipeline information according to the task type and task source of each of the X second acquisition tasks, where the first pipeline information includes the acquisition pipeline type and pipeline parallelism of each of the X second acquisition tasks; determine a first to-be-executed queue according to the first preset start time, the first pipeline information, the current time, the task type priority, the time statistical information, and the estimated consumption duration.

[0042] In another possible design, the processing module is further configured to determine the first planned start time of each of the X second acquisition tasks according to the first preset start time and standard deviation of each of the X second acquisition tasks; determine a first to-be-executed queue according to the first preset start time, the first pipeline information, the first planned start time, the current time, the task type priority, and the estimated consumption duration.

[0043] In another possible design, the first planned start time satisfies:

[0044] T 1,i =H 1,i -k 1 ×S;

[0045] where T 1,i represents the first planned start time of the i-th second acquisition task among the X second acquisition tasks, H 1,i represents the first preset start time of the i-th second acquisition task among the X second acquisition tasks, S represents the standard deviation, k 1 represents an empirical parameter, and i is an integer greater than 0 and less than or equal to X.

[0046] In another possible design, the processing module is further configured to screen the X second acquisition tasks according to the first preset start time, the first pipeline information, and the current time to obtain Y second acquisition tasks, where Y is an integer greater than 1 and less than or equal to X; determine a first to-be-executed queue according to the Y second acquisition tasks, the first pipeline information, the first planned start time, the task type priority, and the estimated consumption duration.

[0047] In another possible design, the processing module is further configured to subtract the current time from the first preset start time of each of the X second acquisition tasks to obtain the waiting duration of each of the X second acquisition tasks; if the first waiting duration among the X waiting durations is less than or equal to the first preset threshold, then use the second acquisition task corresponding to the first waiting duration as one of the Y second acquisition tasks.

[0048] In another possible design, the processing module is further configured to divide the Y second acquisition tasks according to the acquisition pipeline type to obtain at least one subset of tasks to be sorted; sort the at least one subset of tasks to be sorted according to the first planned start time, task type priority, and estimated consumption duration to obtain the first execution queue corresponding to each acquisition pipeline type.

[0049] In another possible design, the third execution information includes the task type, task source, and second preset start time of each of the P third acquisition tasks.

[0050] In another possible design, the processing module is further configured to determine the second pipeline information according to the task type and task source of each of the P third acquisition tasks, where the second pipeline information includes the acquisition pipeline type and pipeline parallelism of each of the P third acquisition tasks; determine the second execution queue according to the second preset start time, second pipeline information, current time, task type priority, time statistical information, and estimated consumption duration.

[0051] In another possible design, the processing module is further configured to determine the second planned start time of each of the P third acquisition tasks according to the second preset start time and standard deviation of each of the P third acquisition tasks; add the estimated consumption duration to the second planned start time of each of the P third acquisition tasks to obtain the planned completion time of each of the P third acquisition tasks; determine the second execution queue according to the second preset start time, second pipeline information, planned completion time, current time, task type priority, and estimated consumption duration.

[0052] In another possible design, the second planned start time satisfies:

[0053] T 2,j =H 2,j -k 1 ×S;

[0054] Wherein, T 2,j represents the second planned start time of the jth third acquisition task among the P third acquisition tasks, and H 2,jrepresents the second preset start time of the j-th third acquisition task among P third acquisition tasks, S represents the standard deviation, and k 1 represents the empirical parameter, and j is an integer greater than 0 and less than or equal to P.

[0055] In another possible design, the processing module is further configured to screen the P third acquisition tasks according to the second preset start time, the second pipeline information, and the current time to obtain Q third acquisition tasks, where Q is an integer greater than 1 and less than or equal to P; determine the second to-be-executed queue according to the Q third acquisition tasks, the second pipeline information, the planned completion time, the task type priority, and the estimated consumption duration.

[0056] In another possible design, the processing module is further configured to subtract the current time from the second preset start time of each third acquisition task among the P third acquisition tasks to obtain the waiting duration of each third acquisition task among the P third acquisition tasks; if the second waiting duration among the P waiting durations is less than or equal to the second preset threshold, then use the third acquisition task corresponding to the second waiting duration as one of the Q third acquisition tasks.

[0057] In another possible design, the processing module is further configured to divide the Q third acquisition tasks according to the acquisition pipeline type to obtain at least one adjacent task subset; sort the at least one adjacent task subset according to the planned completion time, the task type priority, and the estimated consumption duration to obtain the second to-be-executed queue corresponding to each acquisition pipeline type.

[0058] In another possible design, the processing module is further configured to determine the maximum pipeline parallelism of each acquisition pipeline type according to the first pipeline information and the second pipeline information; for the first to-be-executed queue corresponding to each acquisition pipeline type, subtract the pipeline parallelism of the m-th second acquisition task in the first to-be-executed queue from the maximum pipeline parallelism to obtain the free pipeline resources corresponding to each acquisition pipeline type, where m is an integer greater than 0 and less than or equal to Y; determine the data acquisition queue according to the free pipeline resources, the first to-be-executed queue, the second to-be-executed queue, and the second pipeline information.

[0059] In another possible design, the processing module is further configured to, for the second to-be-executed queue corresponding to each acquisition pipeline type, determine whether the pipeline parallelism of the third acquisition task in the second to-be-executed queue is less than or equal to the free pipeline resources; if the first pipeline parallelism of the n n-th third acquisition task in the second to-be-executed queue is less than or equal to the free pipeline resources corresponding to the first pipeline parallelism, then successively arrange the m-th second acquisition task in the first to-be-executed queue and the n-th third acquisition task in the data acquisition queue, where n is an integer greater than 0 and less than or equal to Q.

[0060] The operations performed by the metadata collection device and the beneficial effects can be referred to the method and beneficial effects in the first aspect above, and the repeated parts will not be elaborated.

[0061] In a third aspect, an embodiment of the present application provides a metadata collection system, which includes a flexible collection module, a scheduling plan module, a scheduling execution module, and a data storage module. Among them, the flexible collection module is used to obtain historical execution information, the current time, and the task type priority, determine a data collection queue according to the historical execution information, the current time, and the task type priority, and execute instructions such as feature recognition, task planning, and resource planning in the first aspect; the scheduling plan module is used to write a plan interface based on the data collection queue and generate a task execution plan; the scheduling execution module is used to perform data collection through a collection actuator based on the task execution plan; the data storage module is used to store data.

[0062] In a fourth aspect, an embodiment of the present application provides a metadata collection system, which includes a processor, a memory, and a communication bus. Among them, the memory is used to store computer execution instructions; the processor is used to execute the computer execution instructions stored in the memory so that the metadata collection system executes the method described in any item in the first aspect; the communication bus is used to realize the connection and communication between the processor and the memory.

[0063] In a fifth aspect, an embodiment of the present application provides a metadata collection system, which can execute the method described in the first aspect. The functions of this metadata collection system can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. The system can be software and / or hardware.

[0064] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, which is used to store a computer program. When the computer program is executed, the method described in any item in the first aspect is realized.

[0065] In a seventh aspect, an embodiment of the present application provides a computer program product including a computer program. When the computer program is executed, the method described in any item in the first aspect is realized.

[0066] In an eighth aspect, a chip is provided. The chip includes a processor and a communication interface. The communication interface is used to communicate with external devices or internal devices, and the processor is used to implement the methods in the above aspects.

[0067] In a possible design, the chip may further include a memory which stores computer programs or instructions. The processor is configured to execute the computer programs or instructions stored in the memory, or those derived from other programs or instructions. When the computer programs or instructions are executed, the processor is used to implement the methods in the above aspects.

[0068] In another possible design, the chip may be integrated on a metadata collection system or device. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background art, the following will describe the drawings required to be used in the embodiments of the present application or the background art.

[0070] Figure 1 It is a schematic diagram of data distribution;

[0071] Figure 2 It is a schematic diagram of resource consumption;

[0072] Figure 3 It is another schematic diagram of resource consumption;

[0073] Figure 4 It is a schematic structural diagram of a metadata collection system provided by an embodiment of the present application;

[0074] Figure 5 It is a schematic flowchart of a metadata collection method provided by an embodiment of the present application;

[0075] Figure 6 It is a schematic flowchart of a process for planning a data collection queue provided by an embodiment of the present application;

[0076] Figure 7 It is another schematic flowchart of a process for planning a data collection queue provided by an embodiment of the present application;

[0077] Figure 8 It is an example diagram of source - side ETL task data collection provided by an embodiment of the present application;

[0078] Figure 9 It is a schematic structural diagram of a metadata collection device provided by an embodiment of the present application;

[0079] Figure 10 It is a schematic structural diagram of a collection device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0080] The following explains some of the terms involved in the present application to facilitate understanding by those skilled in the art.

[0081] 1. Metadata, which refers to data about data, mainly describes information about the properties of data. According to different uses, metadata can be classified into three categories: business metadata, technical metadata, and operational metadata.

[0082] 2. Task types are divided into different types according to the different metadata types at the source end. Task types can include database tables, data warehouse technologies (extract-transform-load, ETL) tasks, data models, data services, etc.

[0083] 3. Task sources, that is, the source application systems, abbreviated as the source end, refer to the sources of metadata collection, such as enterprise resource planning (ERP), manufacturing execution system (MES), BDI, etc.

[0084] 4. Pipeline, a management unit abstracted by the data source application end and the metadata collection process together, and system resources are allocated and isolated according to the pipeline.

[0085] 5. Pipeline parallelism, the maximum number of parallel tasks supported by a certain pipeline at the same time.

[0086] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.

[0087] It should be understood that in the description of the present application, "at least one" means one or more, and "multiple" means two or more. In addition, words such as "first" and "second" are only used for the purpose of distinguishing descriptions unless otherwise specified, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.

[0088] The following introduces several metadata collection methods:

[0089] Timed collection method, by designing a certain scheduling frequency, collect metadata of the source application system regularly.

[0090] Real-time collection method, by a very fast scanning frequency, read the metadata of the source application system.

[0091] In the existing technical solutions, both the timed collection method and the real-time collection method call the collector to collect metadata according to a fixed task scheduling plan, without considering data characteristics. As Figure 1 shown, Figure 1It is a schematic diagram of data distribution. The schematic diagram of the data distribution includes a data distribution curve 101. Among them, the horizontal axis represents time, the vertical axis represents probability, and the data distribution curve 101 represents the probability of invalid acquisition corresponding to different times. When metadata acquisition is intensive, there may be invalid acquisition and duplicate acquisition in the timed acquisition method and the real-time acquisition method.

[0092] Moreover, in the existing technical solutions, the timed acquisition method and the real-time acquisition method do not consider the resource occupancy problem either. As Figure 2 shown, Figure 2 It is a schematic diagram of resource consumption. Among them, the horizontal axis represents time, and the vertical axis represents resource consumption; different times correspond to different numbers of acquisition tasks, and the resource consumption of different acquisition tasks is also different; when the number of acquisition tasks is large, it will exceed the resource limit and cause task jams.

[0093] In the existing technical solutions, the timed acquisition method has low real-time performance and low resource consumption; the real-time acquisition method has high real-time performance and high resource consumption. As Figure 3 shown, Figure 3 It is another schematic diagram of resource consumption. Among them, the horizontal axis represents real-time performance, and the vertical axis represents resource consumption. It can be seen that the timed acquisition method cannot meet the real-time requirement of metadata acquisition, and the real-time acquisition method consumes the performance of the CPU.

[0094] To solve the above technical problems, the embodiments of the present application provide the following solutions.

[0095] As Figure 4 shown, Figure 4 It is a schematic diagram of the structure of a metadata acquisition system provided by an embodiment of the present application. The metadata acquisition method provided by the present application is applicable to this metadata acquisition system. The metadata acquisition system includes a flexible acquisition module 401, a scheduling plan module 402, a scheduling execution module 403, and a data storage module 404. Among them, the detailed descriptions of each module are as follows.

[0096] The flexible acquisition module 401 is used to obtain historical execution information, the current time, and the task type priority. The historical execution information includes the first execution information of M first acquisition tasks, the second execution information of X second acquisition tasks, and the third execution information of P third acquisition tasks. The first acquisition tasks are completed tasks, the second acquisition tasks are running tasks, and the third acquisition tasks are unexecuted tasks. M is an integer greater than 1, X is an integer greater than 1, and P is an integer greater than 1. Based on the first execution information, determine the time statistical information and the estimated consumption duration of the M first acquisition tasks. Based on the second execution information, the current time, the task type priority, the time statistical information, and the estimated consumption duration, determine the first pending execution queue. Based on the third execution information, the current time, the task type priority, the time statistical information, and the estimated consumption duration, determine the second pending execution queue. Based on the first pending execution queue, the second pending execution queue, the second execution information, and the third execution information, determine the data acquisition queue.

[0097] Among them, the flexible acquisition module 401 includes a status perception module and a balance planning module. The status perception module is used for data acquisition and feature recognition, and the balance planning module is used for planning pipeline parallelism and generating a data acquisition queue.

[0098] The scheduling plan module 402 is used to generate a task execution plan based on the data acquisition queue and write it to the plan interface.

[0099] Among them, the scheduling plan module 402 includes a resource scheduler, and the task execution plan includes a scheduling plan and a scheduling log.

[0100] The scheduling execution module 403 is used to perform data acquisition through the acquisition executor based on the task execution plan.

[0101] Among them, the scheduling execution module 403 includes an acquisition executor.

[0102] Optionally, the task types of the data acquired through the acquisition executor may include database types, ETL task types, and other types.

[0103] Optionally, the database types may include library basic information acquisition, data table acquisition, and view acquisition.

[0104] Optionally, the ETL task types may include integrated operator acquisition and data processing operator acquisition.

[0105] The data storage module 404 is used to store data.

[0106] Among them, the data storage module 404 includes a metadata repository, and the metadata repository is used to store the data acquired through the scheduling execution module 403.

[0107] In an embodiment of the present application, the data storage module 404 may include volatile memories, such as nonvolatile random access memory (NVRAM), phase change RAM (PRAM), magnetoresistive RAM (MRAM), etc., and may also include nonvolatile memories, such as at least one disk storage device, electrically erasable programmable read-only memory (EEPROM), flash memory devices, such as NOR flash memory or NAND flash memory, semiconductor devices, such as solid state disk (SSD), etc.

[0108] It should be noted that the above metadata collection system may be a system that interacts with users. This system may be a software system, a hardware system, or a system combining software and hardware. The present application does not make specific limitations on this. It should also be noted that Figure 4 only an exemplary structural schematic diagram of the metadata collection system is shown, and in actual applications, corresponding transformations can be made to the Figure 4 metadata collection system according to specific circumstances.

[0109] As Figure 5 shown, Figure 5 is a flowchart of a metadata collection method provided by an embodiment of the present application. This method is applicable to the Figure 4 shown metadata collection system. This method includes but is not limited to the following steps:

[0110] Step S501: Obtain historical execution information, current time, and task type priorities. The historical execution information includes the first execution information of M first collection tasks, the second execution information of X second collection tasks, and the third execution information of P third collection tasks.

[0111] Among them, the first collection task is a completed task, the second collection task is a running task, the third collection task is a non-running task, M is an integer greater than 1, X is an integer greater than 1, and P is an integer greater than 1.

[0112] In one implementation, obtain historical execution information from a database (including a local database or a database of other devices). For example, the first execution information of M first collection tasks, the second execution information of X second collection tasks, and the third execution information of P third collection tasks within the first time period of the source application system can be read.

[0113] Among them, the first time period can be any time period in a day, or any time period in several days, or any time period in a certain period of time. For example, the first time period can be from 9:00 to 14:00 from Monday to Wednesday in a week.

[0114] In another implementation, it is also possible to receive historical execution information uploaded by the user. The metadata collection system can provide a data upload interface, and the data upload interface includes a data upload interface. The user can upload the pre-prepared historical execution information to the metadata collection system by clicking the data upload interface.

[0115] In the embodiment of the present application, it is also necessary to obtain the current time and the task type priority.

[0116] Among them, the task types in the task type priority include, but are not limited to, database tables, data integration JOBs, data processing JOBs, APIs, data ETL tasks, data models, and data services.

[0117] Step S502: Determine the time statistical information and the estimated consumption duration of the M first collection tasks according to the first execution information.

[0118] Among them, the first execution information includes the actual start time and the actual completion time of each of the M first collection tasks, and the time statistical information includes an expected value and a standard deviation.

[0119] Specifically, according to the actual start time and the actual completion time of the M first collection tasks, the expected value and the standard deviation are calculated, and the expected value is determined as the estimated consumption duration. Further, the actual completion time of each of the M first collection tasks is subtracted from the actual start time of each of the M first collection tasks to obtain the actual consumption duration of each of the M first collection tasks. Then, according to the actual consumption duration of each of the M first collection tasks and the actual task quantity M of the first collection tasks, the expected value and the standard deviation are calculated, and then the expected value is determined as the estimated consumption duration.

[0120] Among them, the actual consumption duration satisfies:

[0121] x h = ET h - ST h ;

[0122] Among them, x h represents the actual consumption duration of the h-th first collection task among the M first collection tasks, ET h represents the actual completion time of the h-th first collection task among the M first collection tasks, STh denotes the actual start time of the h-th first acquisition task among the M first acquisition tasks, where h is an integer greater than 0 and less than or equal to M.

[0123] The expected values of the M first acquisition tasks satisfy:

[0124] E = ∑(x h ) / M;

[0125] where E represents the expected value of the M first acquisition tasks, and x h denotes the actual duration consumed by the h-th first acquisition task among the M first acquisition tasks, and M represents the actual number of tasks of the first acquisition task.

[0126] The standard deviation of the M first acquisition tasks satisfies:

[0127]

[0128] where S represents the standard deviation of the M first acquisition tasks, and x h denotes the actual duration consumed by the h-th first acquisition task among the M first acquisition tasks, E represents the expected value of the M first acquisition tasks, and M represents the actual number of tasks of the first acquisition task.

[0129] It should be noted that the first execution information may further include the task name, task type, task source, and preset start time of each of the M first acquisition tasks.

[0130] Step S503: Determine the first task queue to be executed according to the second execution information, the current time, the task type priority, the time statistics information, and the estimated duration.

[0131] where the second execution information includes the task type, task source, and first preset start time of each of the X second acquisition tasks.

[0132] First, determine the first pipeline information according to the task type and task source of each of the X second acquisition tasks. The first pipeline information includes the acquisition pipeline type and pipeline parallelism of each of the X second acquisition tasks.

[0133] Among them, the acquisition pipeline type can be obtained by first manual judgment and then background configuration; the pipeline parallelism can be obtained based on the existing network configuration and manual judgment according to the service requirements.

[0134] For example, as shown in Table 1, based on the task type and task source of each second collection task, the collection pipeline type and pipeline parallelism of each second collection task are configured. For the first row, the task type is database table and the task source is ERP, so the collection pipeline type corresponding to the first row is set to ERP pipeline and the pipeline parallelism is set to 1500; for the second row, the task type is database table and the task source is MES, so the collection pipeline type corresponding to the second row is set to MES pipeline and the pipeline parallelism is set to 1800; for the third row, the task type is data integration JOB and the task source is BDI, so the collection pipeline type corresponding to the third row is set to BDI - data integration pipeline and the pipeline parallelism is set to 1000; for the fourth row, the task type is data model and the task source is BDI, so the collection pipeline type corresponding to the fourth row is set to BDI - data model pipeline and the pipeline parallelism is set to 500. Others are similar and will not be elaborated here. The specific configuration method is not limited in this application.

[0135] Table 1

[0136] Task Type Task Source Collection Pipeline Type Pipeline Parallelism Database Table ERP ERP Pipeline 1500 Database Table MES MES Pipeline 1800 Data Integration JOB BDI BDI - Data Integration Pipeline 1000 Data Model BDI BDI - Data Model Pipeline 500 … … … …

[0137] It should be noted that the second execution information may also include the task name of each second collection task among the X second collection tasks.

[0138] According to the first preset start time and standard deviation of each second collection task among the X second collection tasks, determine the first planned start time of each second collection task among the X second collection tasks.

[0139] Among them, the first planned start time satisfies:

[0140] T 1,i =H 1,i -k 1 ×S;

[0141] Among them, T 1,i represents the first planned start time of the i-th second collection task among the X second collection tasks, H 1,i represents the first preset start time of the i-th second collection task among the X second collection tasks, S represents the standard deviation, and k 1 represents the empirical parameter, and i is an integer greater than 0 and less than or equal to X.

[0142] For example, k 1 can be set to 1, then the first planned start time T 1,i =H 1,i -S.

[0143] Secondly, according to the first preset start time, the first pipeline information, and the current time, screen the X second collection tasks to obtain Y second collection tasks.

[0144] Among them, Y is an integer greater than 1 and less than or equal to X.

[0145] Specifically, subtract the current time from the first preset start time of each second collection task among the X second collection tasks to obtain the waiting duration of each second collection task among the X second collection tasks; if the first waiting duration among the X waiting durations is less than or equal to the first preset threshold, then use the second collection task corresponding to the first waiting duration as one of the Y second collection tasks.

[0146] Among them, the first preset threshold is an empirical parameter.

[0147] For example, the X second collection tasks include Task 1, Task 2, and Task 3. Among them, the waiting duration corresponding to Task 1 is 3 minutes, the waiting duration corresponding to Task 2 is 6 minutes, and the waiting duration corresponding to Task 3 is 7 minutes. The first preset threshold is set to 5 minutes; since 3 minutes is less than 5 minutes, 6 minutes is greater than 5 minutes, and 7 minutes is greater than 5 minutes, that is, Task 1 meets the screening conditions, while Task 2 and Task 3 do not meet the screening conditions, then use Task 1 as one of the Y second collection tasks.

[0148] Then, divide the Y second collection tasks according to the collection pipeline type to obtain at least one subset of tasks to be sorted.

[0149] For example, the Y second collection tasks include Task 1, Task 2, Task 3, and Task 4. Among them, the collection pipeline type of Task 1 is the ERP pipeline, the collection pipeline type of Task 2 is the MES pipeline, the collection pipeline type of Task 3 is the MES pipeline, and the collection pipeline type of Task 4 is the ERP pipeline; divide Task 1, Task 2, Task 3, and Task 4 according to the collection pipeline type, then Task 1 and Task 4 belong to the subset of tasks to be sorted corresponding to the ERP pipeline, and Task 2 and Task 3 belong to the subset of tasks to be sorted corresponding to the MES pipeline.

[0150] It should be noted that for the X - Y second collection tasks with waiting durations greater than the first preset threshold, as the current time changes continuously, the waiting duration of each second collection task among the X - Y second collection tasks gradually decreases. If the third waiting duration among the X - Y waiting durations is less than or equal to the first preset threshold, then the second collection task corresponding to the third waiting duration can be used as one of the Y second collection tasks.

[0151] Finally, according to the first planned start time, the task type priority, and the estimated duration, at least one subset of tasks to be sorted is sorted to obtain the first execution queue corresponding to each type of collection pipeline.

[0152] For example, at least one subset of tasks to be sorted includes the subset of tasks to be sorted corresponding to the ERP pipeline and the subset of tasks to be sorted corresponding to the MES pipeline; according to the first planned start time, the task type priority, and the estimated duration, the second collection tasks in the subset of tasks to be sorted corresponding to the ERP pipeline are sorted to obtain the first execution queue corresponding to the ERP pipeline, and the second collection tasks in the subset of tasks to be sorted corresponding to the MES pipeline are sorted to obtain the first execution queue corresponding to the MES pipeline.

[0153] It should be noted that any one or more of the first planned start time, the task type priority, and the estimated duration can be used as the primary sorting criterion. The specific sorting method is not limited in this application.

[0154] Step S504: Determine the second execution queue according to the third execution information, the current time, the task type priority, the time statistical information, and the estimated duration.

[0155] Among them, the third execution information includes the task type, the task source, and the second preset start time of each of the P third collection tasks.

[0156] First, according to the task type and the task source of each of the P third collection tasks, the second pipeline information is determined. The second pipeline information includes the collection pipeline type and the pipeline parallelism of each of the P third collection tasks. The specific determination process refers to step S503 and will not be elaborated here.

[0157] It should be noted that the third execution information may also include the task name of each of the P third collection tasks.

[0158] According to the second preset start time and the standard deviation of each of the P third collection tasks, the second planned start time of each of the P third collection tasks is determined, and then the estimated duration is added to the second planned start time of each of the P third collection tasks to obtain the planned completion time of each of the P third collection tasks.

[0159] Among them, the second planned start time satisfies:

[0160] T 2,j =H 2,j -k 1 ×S;

[0161] Among them, T2,j represents the second planned start time of the j-th third acquisition task among P third acquisition tasks, H 2,j represents the second preset start time of the j-th third acquisition task among P third acquisition tasks, S represents the standard deviation, k 1 represents the empirical parameter, and j is an integer greater than 0 and less than or equal to P.

[0162] The planned completion time satisfies:

[0163] T 3,j = T 2,j + R;

[0164] R = E;

[0165] where, T 3,j represents the planned completion time of the j-th third acquisition task among P third acquisition tasks, T 2,j represents the second planned start time of the j-th third acquisition task among P third acquisition tasks, R represents the expected consumption duration, E represents the expected value, and j is an integer greater than 0 and less than or equal to P.

[0166] Secondly, according to the second preset start time, the second pipeline information, and the current time, screen the P third acquisition tasks to obtain Q third acquisition tasks.

[0167] where, Q is an integer greater than 1 and less than or equal to P.

[0168] Specifically, subtract the current time from the second preset start time of each third acquisition task among the P third acquisition tasks to obtain the waiting duration of each third acquisition task among the P third acquisition tasks; if the second waiting duration among the P waiting durations is less than or equal to the second preset threshold, then use the third acquisition task corresponding to the second waiting duration as one of the Q third acquisition tasks.

[0169] where, the second preset threshold is an empirical parameter.

[0170] For example, the P third acquisition tasks include task A, task B, and task C. Among them, the waiting duration corresponding to task A is 2 minutes, the waiting duration corresponding to task B is 1 minute, and the waiting duration corresponding to task C is 5 minutes. The second preset threshold is set to 1 minute; since 2 minutes is greater than 1 minute, 1 minute is less than or equal to 1 minute, and 5 minutes is greater than 1 minute, that is, task B meets the screening conditions, while task A and task C do not meet the screening conditions, and task B is used as one of the Q third acquisition tasks.

[0171] It should be noted that for the P-Q third collection tasks with waiting durations greater than the second preset threshold, as the current time continuously changes, the waiting duration of each of the P-Q third collection tasks gradually decreases. If the fourth waiting duration among the P-Q waiting durations is less than or equal to the second preset threshold, the third collection task corresponding to the fourth waiting duration can be used as one of the Q third collection tasks.

[0172] Then, according to the collection pipeline type, the Q third collection tasks are divided to obtain at least one adjacent task subset.

[0173] For example, the Q third collection tasks include task A, task B, task C, and task D. Among them, the collection pipeline type of task A is the ERP pipeline, the collection pipeline type of task B is the MES pipeline, the collection pipeline type of task C is the MES pipeline, and the collection pipeline type of task D is the BDI - data integration pipeline; according to the collection pipeline type, tasks A, B, C, and D are divided, then task A belongs to the adjacent task subset corresponding to the ERP pipeline, tasks B and C belong to the adjacent task subset corresponding to the MES pipeline, and task D belongs to the adjacent task subset corresponding to the BDI - data integration pipeline.

[0174] Finally, according to the planned completion time, task type priority, and estimated consumption duration, the at least one adjacent task subset is sorted to obtain a second to-be-executed queue corresponding to each collection pipeline type.

[0175] For example, the at least one adjacent task subset includes the adjacent task subset corresponding to the ERP pipeline, the adjacent task subset corresponding to the MES pipeline, and the adjacent task subset corresponding to the BDI - data integration pipeline; according to the planned completion time, task type priority, and estimated consumption duration, the third collection tasks in the adjacent task subset corresponding to the ERP pipeline are sorted to obtain the second to-be-executed queue corresponding to the ERP pipeline, and the third collection tasks in the adjacent task subset corresponding to the MES pipeline are sorted to obtain the second to-be-executed queue corresponding to the MES pipeline, and the third collection tasks in the adjacent task subset corresponding to the BDI - data integration pipeline are sorted to obtain the second to-be-executed queue corresponding to the BDI - data integration pipeline.

[0176] It should be noted that any one or more of the planned completion time, task type priority, and estimated consumption duration can be used as the primary index for sorting. The specific sorting method is not limited in this application.

[0177] Step S505: Determine a data collection queue according to the first to-be-executed queue, the second to-be-executed queue, the second execution information, and the third execution information.

[0178] First, based on the first pipeline information and the second pipeline information, determine the maximum pipeline parallelism for each type of acquisition pipeline.

[0179] Among them, the maximum pipeline parallelism for each type of acquisition pipeline is the pipeline resource constraint of the data acquisition queue.

[0180] Specifically, based on the pipeline information of Y second acquisition tasks and the pipeline information of Q third acquisition tasks, determine the types of acquisition pipelines in the Y second acquisition tasks and the Q third acquisition tasks, as well as the maximum pipeline parallelism for each type of acquisition pipeline.

[0181] For example, the Y second acquisition tasks include Task 1, Task 2, Task 3, and Task 4, and the Q third acquisition tasks include Task 5, Task 6, Task 7, and Task 8. As shown in Table 2, based on the acquisition pipeline types and pipeline parallelisms of Tasks 1 to 8, determine the maximum pipeline parallelism for each type of acquisition pipeline. The acquisition pipeline type corresponding to Task 1 in the first row is the ERP pipeline, and the pipeline parallelism is 1500. The acquisition pipeline type corresponding to Task 2 in the second row is the MES pipeline, and the pipeline parallelism is 1800. The acquisition pipeline type corresponding to Task 3 in the third row is the BDI - data integration pipeline, and the pipeline parallelism is 1400. The acquisition pipeline type corresponding to Task 4 in the fourth row is the ERP pipeline, and the pipeline parallelism is 1300. The acquisition pipeline type corresponding to Task 5 in the fifth row is the ERP pipeline, and the pipeline parallelism is 1600. The acquisition pipeline type corresponding to Task 6 in the sixth row is the ERP pipeline, and the pipeline parallelism is 1300. The acquisition pipeline type corresponding to Task 7 in the seventh row is the MES pipeline, and the pipeline parallelism is 1000. The acquisition pipeline type corresponding to Task 8 in the eighth row is the BDI - data model pipeline, and the pipeline parallelism is 800. Based on the above pipeline information, it can be determined that the types of acquisition pipelines include the ERP pipeline, the MES pipeline, the BDI - data integration pipeline, and the BDI - data model pipeline, and the maximum pipeline parallelism of the ERP pipeline is 1600, the maximum pipeline parallelism of the MES pipeline is 1800, the maximum pipeline parallelism of the BDI - data integration pipeline is 1400, and the maximum pipeline parallelism of the BDI - data model pipeline is 800.

[0182] Table 2

[0183]

[0184] Then, for the first execution queue corresponding to each type of acquisition pipeline, subtract the pipeline parallelism of the m - th second acquisition task in the first execution queue from the maximum pipeline parallelism to obtain the remaining pipeline resources corresponding to each type of acquisition pipeline.

[0185] Among them, m is an integer greater than 0 and less than or equal to Y.

[0186] Finally, determine the data collection queue according to the available pipeline resources, the first to-be-executed queue, the second to-be-executed queue, and the second pipeline information.

[0187] Specifically, for the second to-be-executed queue corresponding to each collection pipeline type, determine whether the pipeline parallelism of the third collection task in the second to-be-executed queue is less than or equal to the available pipeline resources; if the first pipeline parallelism of the nth third collection task in the second to-be-executed queue is less than or equal to the available pipeline resources corresponding to the first pipeline parallelism, then arrange the mth second collection task and the nth third collection task in the first to-be-executed queue into the data collection queue in sequence; if the first pipeline parallelism of the nth third collection task in the second to-be-executed queue is greater than the available pipeline resources corresponding to the first pipeline parallelism, then continue to determine whether the pipeline parallelism of the (n + 1)th third collection task in the second to-be-executed queue is less than or equal to the available pipeline resources.

[0188] Wherein, n is an integer greater than 0 and less than or equal to Q.

[0189] For example, the collection pipeline type includes the ERP pipeline; arrange the mth second collection task in the first to-be-executed queue corresponding to the ERP pipeline into the data collection queue corresponding to the ERP pipeline, then subtract the pipeline parallelism of the mth second collection task from the maximum pipeline parallelism of the ERP pipeline to obtain the available pipeline resources corresponding to the ERP pipeline; then sequentially determine whether the pipeline parallelism of the third collection task in the second to-be-executed queue corresponding to the ERP pipeline is less than or equal to the available pipeline resources corresponding to the ERP pipeline. If the first pipeline parallelism of the nth third collection task in the second to-be-executed queue is less than or equal to the available pipeline resources corresponding to the ERP pipeline, then arrange the nth third collection task in the second to-be-executed queue into the data collection queue corresponding to the ERP pipeline.

[0190] For example, as Figure 6 shown, Figure 6 FIG. is a schematic flow chart of planning a data collection queue provided by an embodiment of the present application. Arrange the collection tasks in the first to-be-executed queue and the collection tasks in the second to-be-executed queue into the pipeline resources, and then determine whether there are available pipeline resources. If there are available pipeline resources, it means that the pipeline resource constraint is exceeded, and arrange the collection tasks in the first to-be-executed queue and the collection tasks in the second to-be-executed queue into the data collection queue; if there are no available pipeline resources, it means that the pipeline resource constraint is exceeded, and it is necessary to rescan the second to-be-executed queue.

[0191] For example, as Figure 7 shown, Figure 7It is another schematic flowchart of planning a data collection queue provided by an embodiment of the present application. After a task is started, for each type of collection pipeline, a to-be-sorted list in the collection pipeline is identified. The to-be-sorted list includes pipeline resource constraints, a subset of tasks to be sorted, and a subset of adjacent tasks; the subset of tasks to be sorted is sorted to obtain a first to-be-executed queue, and the subset of adjacent tasks is sorted to obtain a second to-be-executed queue; for the pipeline resource constraints, the tasks in the first to-be-executed queue are arranged into the pipeline resources, and then the TOP task in the second to-be-executed queue is taken out. If the TOP task in the second to-be-executed queue is empty, the operation ends; if the TOP task in the second to-be-executed queue is not empty, it is determined whether the pipeline resource constraints are exceeded. If the pipeline resource constraints are not exceeded, the tasks in the first to-be-executed queue and the TOP task in the second to-be-executed queue are added to the data collection queue; if the pipeline resource constraints are exceeded, the tasks in the second to-be-executed queue are cyclically obtained and the idle pipeline resources are filled.

[0192] Step S506: Perform data collection based on the data collection queue.

[0193] Specifically, based on the data collection queue written into the planning interface, the resource scheduler generates a task execution plan; when the planned start time is met and the collection executor matches the source application system, data collection is performed through the collection executor based on the task execution plan.

[0194] In the embodiment of the present application, after a task is completed and executed, it dequeues. If the task execution fails, the task will regenerate a to-be-sorted list.

[0195] Optionally, if the source-side JOB is still running after the tasks in the collection pipeline are completed, a new task is started, and the planned start times of other tasks are postponed by k 2 ×S minutes until the source-side JOB ends.

[0196] Among them, k 2 represents an empirical parameter.

[0197] It should be noted that the planned start time can be postponed once or multiple times, and the present application does not make any restrictions.

[0198] For example, the task collection list includes tasks 1, 2, 3, and 4, and k 2 is set to 0.5. If the source-side JOB is still running after task 1 in the collection pipeline is completed, task 2 is started, and the planned start times of tasks 3 and 4 are postponed by 0.5×S minutes; if the source-side JOB is still running after task 2 is completed, task 3 is started, and the planned start time of task 4 is postponed by 0.5×S minutes again.

[0199] For example, as Figure 8 shownFigure 8 It is an example diagram of source - end ETL task data collection provided by an embodiment of this application. Among them, the first row represents the collection situation of ODS - Task A, the second row represents the collection situation of DWI - Task B, the third row represents the collection situation of DWR - Task C, and the fourth row represents the collection situation of DWR - Task D.

[0200] Specifically, Figure 8 in Figure 8 , "ST" represents the actual start time of the task; "ET" represents the actual completion time of the task; "Day1" and "Day2" represent historical collection times; "Day3 - TOBE" represents the future collection time; ODS - Task A includes A Task - Instance 1, A Task - Instance 2, and A Task - Instance 3; DWI - Task B includes B Task - Instance 1, B Task - Instance 2, and Data Task 1; DWR - Task C includes C Task - Instance 1, C Task - Instance 2, C Task - Instance 3, C Task - Instance 4, and Data Task 2; DWR - Task D includes D Task - Instance 1, D Task - Instance 2, and Data Task 3.

[0201] In Figure 8 the example, the historical execution information includes the data collected on Day1 and Day2. Among them, on Day1, the data collection of A Task - Instance 1, B Task - Instance 1, C Task - Instance 1, C Task - Instance 2, and D Task - Instance 1 has been completed, and on Day2, the data collection of A Task - Instance 2, B Task - Instance 2, C Task - Instance 3, C Task - Instance 4, and D Task - Instance 2 has been completed; in this example, the source - end data statistical characteristics can be identified based on the historical execution information, and then the planned start time 1 of the collection task for Day3 - TOBE can be determined, and the data collection queues for A Task - Instance 3, Data Task 1, Data Task 2, and Data Task 3 can be planned.

[0202] In the embodiment of this application, by combining the historical execution information, the current time, and the task - type priority for queue planning, the first to - be - executed queue and the second to - be - executed queue are obtained. Then, the pipeline resource limit is determined through the second execution information and the third execution information. Then, based on the pipeline resource limit, the first to - be - executed queue and the second to - be - executed queue are planned, and then the data collection queue is determined. Through the data collection queue, accurate collection can be performed based on data characteristics, and by planning the number of collection tasks through the pipeline resource limit, it is possible to avoid system suspension and blockage when the quantity is large, reduce the resource consumption of real - time data collection, and improve the collection efficiency and system stability.

[0203] As Figure 9 shown, Figure 9It is a schematic structural diagram of a metadata collection device provided by an embodiment of the present application. Among them, the metadata collection device includes an acquisition module 901 and a processing module 902. The detailed descriptions of each unit are as follows.

[0204] Acquisition module 901: It is used to acquire historical execution information, current time, and task type priorities. The historical execution information includes the first execution information of M first collection tasks, the second execution information of X second collection tasks, and the third execution information of P third collection tasks. The first collection tasks are completed tasks, the second collection tasks are in-running tasks, and the third collection tasks are un-run tasks. M is an integer greater than 1, X is an integer greater than 1, and P is an integer greater than 1.

[0205] Processing module 902: It is used to determine the time statistical information and estimated consumption duration of M first collection tasks according to the first execution information.

[0206] Optionally, the first execution information includes the actual start time and actual completion time of each of the M first collection tasks.

[0207] Optionally, the processing module 902 is further used to calculate the time statistical information according to the actual start time and actual completion time. The time statistical information includes the expected value and standard deviation; and determine the expected consumption duration as the expected value.

[0208] Processing module 902: It is further used to determine the first pending execution queue according to the second execution information, current time, task type priorities, time statistical information, and estimated consumption duration.

[0209] Optionally, the second execution information includes the task type, task source, and first preset start time of each of the X second collection tasks.

[0210] Optionally, the processing module 902 is further used to determine the first pipeline information according to the task type and task source of each of the X second collection tasks. The first pipeline information includes the collection pipeline type and pipeline parallelism of each of the X second collection tasks; and determine the first pending execution queue according to the first preset start time, first pipeline information, current time, task type priorities, time statistical information, and estimated consumption duration.

[0211] Optionally, the processing module 902 is further used to determine the first planned start time of each of the X second collection tasks according to the first preset start time and standard deviation of each of the X second collection tasks; and determine the first pending execution queue according to the first preset start time, first pipeline information, first planned start time, current time, task type priorities, and estimated consumption duration.

[0212] Optionally, the first planned start time satisfies:

[0213] T 1,i = H 1,i - k 1 × S;

[0214] Wherein, T 1,i represents the first planned start time of the i-th second acquisition task among the X second acquisition tasks, H 1,i represents the first preset start time of the i-th second acquisition task among the X second acquisition tasks, S represents the standard deviation, and k 1 represents an empirical parameter, and i is an integer greater than 0 and less than or equal to X.

[0215] Optionally, the processing module 902 is further configured to screen the X second acquisition tasks according to the first preset start time, the first pipeline information, and the current time to obtain Y second acquisition tasks, where Y is an integer greater than 1 and less than or equal to X; determine the first to-be-executed queue according to the Y second acquisition tasks, the first pipeline information, the first planned start time, the task type priority, and the estimated consumption duration.

[0216] Optionally, the processing module 902 is further configured to subtract the current time from the first preset start time of each second acquisition task among the X second acquisition tasks to obtain the waiting duration of each second acquisition task among the X second acquisition tasks; if the first waiting duration among the X waiting durations is less than or equal to the first preset threshold, then use the second acquisition task corresponding to the first waiting duration as one of the Y second acquisition tasks.

[0217] Optionally, the processing module 902 is further configured to divide the Y second acquisition tasks according to the acquisition pipeline type to obtain at least one to-be-sorted task subset; sort the at least one to-be-sorted task subset according to the first planned start time, the task type priority, and the estimated consumption duration to obtain the first to-be-executed queue corresponding to each acquisition pipeline type.

[0218] The processing module 902: is further configured to determine the second to-be-executed queue according to the third execution information, the current time, the task type priority, the time statistical information, and the estimated consumption duration.

[0219] Optionally, the third execution information includes the task type, the task source, and the second preset start time of each third acquisition task among the P third acquisition tasks.

[0220] Optionally, the processing module 902 is further configured to determine second pipeline information according to the task type and task source of each of the P third collection tasks, where the second pipeline information includes the collection pipeline type and pipeline parallelism of each of the P third collection tasks; determine a second to-be-executed queue according to the second preset start time, the second pipeline information, the current time, the task type priority, the time statistical information, and the estimated consumption duration.

[0221] Optionally, the processing module 902 is further configured to determine the second planned start time of each of the P third collection tasks according to the second preset start time and standard deviation of each of the P third collection tasks; add the estimated consumption duration to the second planned start time of each of the P third collection tasks to obtain the planned completion time of each of the P third collection tasks; determine a second to-be-executed queue according to the second preset start time, the second pipeline information, the planned completion time, the current time, the task type priority, and the estimated consumption duration.

[0222] Optionally, the second planned start time satisfies:

[0223] T 2,j = H 2,j - k 1 × S;

[0224] Wherein, T 2,j represents the second planned start time of the jth third collection task among the P third collection tasks, H 2,j represents the second preset start time of the jth third collection task among the P third collection tasks, S represents the standard deviation, and k 1 represents an empirical parameter, and j is an integer greater than 0 and less than or equal to P.

[0225] Optionally, the processing module 902 is further configured to screen the P third collection tasks according to the second preset start time, the second pipeline information, and the current time to obtain Q third collection tasks, where Q is an integer greater than 1 and less than or equal to P; determine a second to-be-executed queue according to the Q third collection tasks, the second pipeline information, the planned completion time, the task type priority, and the estimated consumption duration.

[0226] Optionally, the processing module 902 is further configured to subtract the current time from the second preset start time of each of the P third collection tasks to obtain the waiting duration of each of the P third collection tasks; if the second waiting duration among the P waiting durations is less than or equal to the second preset threshold, use the third collection task corresponding to the second waiting duration as one of the Q third collection tasks.

[0227] Optionally, the processing module 902 is further configured to divide the Q third acquisition tasks according to the acquisition pipeline type to obtain at least one adjacent task subset; and sort the at least one adjacent task subset according to the planned completion time, task type priority, and estimated consumption duration to obtain a second to-be-executed queue corresponding to each acquisition pipeline type.

[0228] The processing module 902: is further configured to determine a data acquisition queue according to the first to-be-executed queue, the second to-be-executed queue, the second execution information, and the third execution information.

[0229] Optionally, the processing module 902 is further configured to determine the maximum pipeline parallelism of each acquisition pipeline type according to the first pipeline information and the second pipeline information; for the first to-be-executed queue corresponding to each acquisition pipeline type, subtract the pipeline parallelism of the m-th second acquisition task in the first to-be-executed queue from the maximum pipeline parallelism to obtain the available pipeline resources corresponding to each acquisition pipeline type, where m is an integer greater than 0 and less than or equal to Y; and determine a data acquisition queue according to the available pipeline resources, the first to-be-executed queue, the second to-be-executed queue, and the second pipeline information.

[0230] Optionally, the processing module 902 is further configured to determine whether the pipeline parallelism of the third acquisition task in the second to-be-executed queue corresponding to each acquisition pipeline type is less than or equal to the available pipeline resources; if the first pipeline parallelism of the n-th third acquisition task in the second to-be-executed queue is less than or equal to the available pipeline resources corresponding to the first pipeline parallelism, then arrange the m-th second acquisition task and the n-th third acquisition task in the first to-be-executed queue into the data acquisition queue in sequence, where n is an integer greater than 0 and less than or equal to Q.

[0231] The processing module 902: is further configured to perform data acquisition based on the data acquisition queue.

[0232] It should be noted that the implementation of each module can also correspond to the corresponding description in the Figure 5 method embodiment shown, and execute the methods and functions performed by the metadata acquisition system in the above embodiments.

[0233] The foregoing content has introduced in detail the metadata acquisition system provided by the present application and the process of generating a data acquisition queue using this system. Next, the Figure 10 deployment method of the metadata acquisition system will be introduced.

[0234] As Figure 10 shown, Figure 10It is a schematic structural diagram of a collection device provided by an embodiment of the present application. The collection device includes a processor 1001, a transceiver 1002, and a memory 1003. Among them, the processor 1001, the transceiver 1002, and the memory 1003 can communicate with each other through a communication bus 1004 connection path to transmit instructions and / or data signals. The memory 1003 is used to store a computer program, and the processor 1001 is used to call and run the computer program from the memory 1003 to control the transceiver 1002 to transmit and receive signals.

[0235] The above-mentioned processor 1001 can correspond to Figure 9 the processing module 902 therein. The above-mentioned processor 1001 and the memory 1003 can be combined into a processing device. The processor 1001 is used to execute the program code stored in the memory 1003 to implement the above functions. Specifically, in implementation, the memory 1003 can also be integrated in the processor 1001 or independent of the processor 1001.

[0236] The above-mentioned transceiver 1002 can also be referred to as a transceiver unit or a transceiver module. The transceiver 1002 can include a receiver (or called a receiver, receiving circuit) and a transmitter (or called a transmitter, transmitting circuit). Among them, the receiver is used to receive signals, and the transmitter is used to send signals.

[0237] It should be understood that Figure 10 the collection device shown can implement Figure 5 each process related to the metadata collection system in the method embodiment shown. The operations and / or functions of each module in the collection device are respectively for implementing the corresponding processes in the above method embodiment. For details, reference can be made to the description in the above method embodiment. To avoid repetition, the detailed description is appropriately omitted here.

[0238] Among them, the processor 1001 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary modules described in combination with the disclosure of the present application. The processor 1001 can also be a combination that implements a computing function, such as a combination including one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The communication bus 1004 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 10It is represented only by a thick line, but it does not mean that there is only one bus or one type of bus. The communication bus 1004 is used to implement the connection and communication between these components. Among them, in the embodiment of the present application, the memory 1003 can be various types of memories mentioned above. The transceiver 1002 is used to communicate instructions or data with other components. The processor can cooperate with the memory and the transceiver to execute any method and function of the metadata acquisition system in the above-mentioned embodiment of the application.

[0239] Optionally, the memory 1003 can also be at least one storage device located far from the aforementioned processor 1001.

[0240] Optionally, a set of computer program codes or configuration information can also be stored in the memory 1003.

[0241] Optionally, the processor 1001 can also execute the program stored in the memory 1003.

[0242] According to the method provided by the embodiment of the present application, the present application also provides a chip system, which includes a processor for supporting the metadata acquisition system or the metadata acquisition device to implement the functions involved in any of the above embodiments, such as generating a data acquisition queue or processing the data characteristics involved in the above method.

[0243] According to the method provided by the embodiment of the present application, the present application also provides a computer program product, which includes: a computer program, when the computer program runs on a computer, it causes the computer to execute Figure 4 or Figure 5 the method of any one of the embodiments shown.

[0244] According to the method provided by the embodiment of the present application, the present application also provides a computer-readable medium, which stores a computer program, when the computer program runs on a computer, it causes the computer to execute Figure 4 or Figure 5 the method of any one of the embodiments shown.

[0245] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The readable medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., high-density digital video disc (DVD)), or a semiconductor medium (e.g., solid state disc (SSD)), etc.

[0246] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone.

[0247] It should be understood that in the embodiments of the present application, "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B solely based on A, and B can also be determined according to A and / or other information.

[0248] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included within the protection scope of the present application.

Claims

1. A metadata collection method, characterized in that, it includes: Obtain historical execution information, current time, and task type priorities. The historical execution information includes the first execution information of M first collection tasks, the second execution information of X second collection tasks, and the third execution information of P third collection tasks. The first collection tasks are completed tasks, the second collection tasks are in-running tasks, and the third collection tasks are unrun tasks. M is an integer greater than 1, X is an integer greater than 1, and P is an integer greater than 1; According to the first execution information, determine the time statistical information and the estimated consumption duration of the M first collection tasks; According to the second execution information, the current time, the task type priorities, the time statistical information, and the estimated consumption duration, determine the first pending execution queue; According to the third execution information, the current time, the task type priorities, the time statistical information, and the estimated consumption duration, determine the second pending execution queue; According to the first pending execution queue, the second pending execution queue, the second execution information, and the third execution information, determine the data collection queue; Based on the data collection queue, perform data collection.

2. The method according to claim 1, characterized in that, the first execution information includes the actual start time and the actual completion time of each of the M first collection tasks; the determining the time statistical information and the estimated consumption duration of the M first collection tasks according to the first execution information includes: According to the actual start time and the actual completion time, calculate the time statistical information, and the time statistical information includes an expected value and a standard deviation; Determine the expected value as the estimated consumption duration.

3. The method according to claim 2, characterized in that, the second execution information includes the task type, task source, and first preset start time of each of the X second collection tasks; the determining the first pending execution queue according to the second execution information, the current time, the task type priorities, the time statistical information, and the estimated consumption duration includes: According to the task type and the task source of each of the X second collection tasks, determine the first pipeline information, and the first pipeline information includes the collection pipeline type and pipeline parallelism of each of the X second collection tasks; According to the first preset start time, the first pipeline information, the current time, the task type priorities, the time statistical information, and the estimated consumption duration, determine the first pending execution queue.

4. The method according to claim 3, characterized in that, the determining the first pending execution queue according to the first preset start time, the first pipeline information, the current time, the task type priorities, the time statistical information, and the estimated consumption duration includes: Determine the first planned start time for each of the X second collection tasks according to the first preset start time and the standard deviation of each of the X second collection tasks; Determine the first queue to be executed according to the first preset start time, the first pipeline information, the first planned start time, the current time, the task type priority, and the estimated duration.

5. The method according to claim 4, characterized in that the first planned start time satisfies: T 1,i = H 1,i - k 1 × S; Among them, the T 1,i represents the first planned start time of the i-th second acquisition task among the X second acquisition tasks, the H 1,i represents the first preset start time of the i-th second acquisition task among the X second acquisition tasks, the S represents the standard deviation, and the k 1 represents an empirical parameter, and i is an integer greater than 0 and less than or equal to X.

6. The method according to claim 4, characterized in that the determining the first queue to be executed according to the first preset start time, the first pipeline information, the first planned start time, the current time, the task type priority, and the estimated duration includes: Screen the X second collection tasks according to the first preset start time, the first pipeline information, and the current time to obtain Y second collection tasks, where Y is an integer greater than 1 and less than or equal to X; Determine the first queue to be executed according to the Y second collection tasks, the first pipeline information, the first planned start time, the task type priority, and the estimated duration.

7. The method according to claim 6, characterized in that the screening the X second collection tasks according to the first preset start time, the first pipeline information, and the current time to obtain Y second collection tasks includes: Subtract the current time from the first preset start time of each of the X second collection tasks to obtain the waiting duration of each of the X second collection tasks; If the first waiting duration among the X waiting durations is less than or equal to the first preset threshold, then use the second collection task corresponding to the first waiting duration as one of the Y second collection tasks.

8. The method according to claim 6, characterized in that the determining the first queue to be executed according to the Y second collection tasks, the first pipeline information, the first planned start time, the task type priority, and the estimated duration includes: Divide the Y second collection tasks according to the collection pipeline type to obtain at least one subset of tasks to be sorted; Sort the at least one subset of tasks to be sorted according to the first planned start time, the task type priority, and the estimated duration to obtain the first queue to be executed corresponding to each collection pipeline type.

9. The method according to any one of claims 2-8, characterized in that the third execution information includes the task type, task source, and second preset start time of each of the P third collection tasks; the determining the second queue to be executed according to the third execution information, the current time, the task type priority, the time statistical information, and the estimated duration includes: Determine second pipeline information according to the task type and task source of each of the P third collection tasks, where the second pipeline information includes the collection pipeline type and pipeline parallelism of each of the P third collection tasks; Determine the second to-be-executed queue according to the second preset start time, the second pipeline information, the current time, the task type priority, the time statistics information, and the estimated consumption duration.

10. The method according to claim 9, characterized in that the determining the second to-be-executed queue according to the second preset start time, the second pipeline information, the current time, the task type priority, the time statistics information, and the estimated consumption duration includes: Determine the second planned start time of each of the P third collection tasks according to the second preset start time and the standard deviation of each of the P third collection tasks; Add the estimated consumption duration to the second planned start time of each of the P third collection tasks to obtain the planned completion time of each of the P third collection tasks; Determine the second to-be-executed queue according to the second preset start time, the second pipeline information, the planned completion time, the current time, the task type priority, and the estimated consumption duration.

11. The method according to claim 10, characterized in that the second planned start time satisfies: T 2,j = H 2,j - k 1 × S; Wherein, the T 2,j represents the second planned start time of the j-th third acquisition task among the P third acquisition tasks, the H 2,j represents the second preset start time of the j-th third acquisition task among the P third acquisition tasks, the S represents the standard deviation, and the k 1 represents an empirical parameter, and j is an integer greater than 0 and less than or equal to P.

12. The method according to claim 10, characterized in that the determining the second to-be-executed queue according to the second preset start time, the second pipeline information, the planned completion time, the current time, the task type priority, and the estimated consumption duration includes: Screen the P third collection tasks according to the second preset start time, the second pipeline information, and the current time to obtain Q third collection tasks, where Q is an integer greater than 1 and less than or equal to P; Determine the second to-be-executed queue according to the Q third collection tasks, the second pipeline information, the planned completion time, the task type priority, and the estimated consumption duration.

13. The method according to claim 12, characterized in that the screening the P third collection tasks according to the second preset start time, the second pipeline information, and the current time to obtain Q third collection tasks includes: Subtract the current time from the second preset start time of each of the P third collection tasks to obtain the waiting duration of each of the P third collection tasks; If the second waiting duration among the P waiting durations is less than or equal to a second preset threshold, then use the third collection task corresponding to the second waiting duration as one of the Q third collection tasks.

14. The method according to claim 12, characterized in that Determining the second to-be-executed queue according to the Q third acquisition tasks, the second pipeline information, the planned completion time, the task type priority, and the estimated consumption duration includes: Dividing the Q third acquisition tasks according to the acquisition pipeline type to obtain at least one adjacent task subset; Sorting the at least one adjacent task subset according to the planned completion time, the task type priority, and the estimated consumption duration to obtain a second to-be-executed queue corresponding to each acquisition pipeline type.

15. The method according to any one of claims 3-14, wherein, Determining the data acquisition queue according to the first to-be-executed queue, the second to-be-executed queue, the second execution information, and the third execution information includes: Determining the maximum pipeline parallelism of each acquisition pipeline type according to the first pipeline information and the second pipeline information; For the first to-be-executed queue corresponding to each acquisition pipeline type, subtracting the pipeline parallelism of the m-th second acquisition task in the first to-be-executed queue from the maximum pipeline parallelism to obtain the available pipeline resources corresponding to each acquisition pipeline type, where m is an integer greater than 0 and less than or equal to Y; Determining the data acquisition queue according to the available pipeline resources, the first to-be-executed queue, the second to-be-executed queue, and the second pipeline information.

16. The method according to claim 15, wherein, Determining the data acquisition queue according to the available pipeline resources, the first to-be-executed queue, the second to-be-executed queue, and the second pipeline information includes: For the second to-be-executed queue corresponding to each acquisition pipeline type, determining whether the pipeline parallelism of the third acquisition task in the second to-be-executed queue is less than or equal to the available pipeline resources; If the first pipeline parallelism of the n-th third acquisition task in the second to-be-executed queue is less than or equal to the available pipeline resources corresponding to the first pipeline parallelism, then successively arranging the m-th second acquisition task and the n-th third acquisition task in the first to-be-executed queue into the data acquisition queue, where n is an integer greater than 0 and less than or equal to Q.

17. A metadata acquisition device, wherein, The metadata acquisition device includes a processor and a memory. The memory is used to store a computer program, and the processor is used to call the computer program to execute the method according to any one of claims 1-16.

18. A computer-readable storage medium, wherein, It is used to store a computer program, and when the computer program runs on a computer, it causes the computer to execute the method according to any one of claims 1-16.

19. A chip, wherein, The chip includes a processor and a communication interface. The communication interface is used to communicate with external devices or internal devices, and the processor is used to implement the method according to any one of claims 1-16.

Citation Information

Cited By

  • Data acquisition system and data acquisition method

    CN122470321A