Data acquisition scheduling method and device, equipment, storage medium and product

By building a dynamic parameter configuration engine and unified interface design, configuration files and scheduling rules are automatically generated, which solves the high cost and heterogeneity problems of self-developed data collection solutions and realizes efficient and reliable data collection and cross-environment scheduling.

CN120705205APending Publication Date: 2025-09-26INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510882437.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Self-developed data collection solutions have a long development cycle and high maintenance costs. The standardized collection logic of third-party services conflicts greatly with the heterogeneity of local business data, resulting in high communication costs and easily broken data links.

Method used

Build a dynamic parameter configuration engine to automatically generate configuration files and scheduling rules, encapsulate third-party service scheduling interfaces into standardized services, and execute data collection and scheduling tasks through a unified interface to reduce manual intervention and data link disruption during environment migration.

Benefits of technology

It improves the adaptability and deployment efficiency of data collection strategies in different environments, reduces communication costs, enhances the efficiency of data transmission and the maintainability of the system, and reduces the risk of failure caused by environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705205A_ABST
    Figure CN120705205A_ABST
Patent Text Reader

Abstract

The invention provides a data acquisition scheduling method and device, equipment, a storage medium and a product, and relates to the field of big data. The method comprises the following steps: acquiring a scheduling parameter of a to-be-scheduled task, determining a corresponding configuration file based on the scheduling parameter, and generating a first data scheduling rule according to the configuration file; calling a data acquisition scheduling interface, and executing a data acquisition scheduling task according to a first data scheduling rule; and under the condition that the data acquisition scheduling task is successfully executed, obtaining scheduling data corresponding to the to-be-scheduled task. According to the method, the dynamic parameter configuration engine is constructed, and the configuration file is automatically generated according to the configuration parameters, so that the automation and intelligence level of data acquisition and scheduling is improved, the manual intervention cost is reduced, and the data scheduling efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of big data, and in particular to a data acquisition and scheduling method, device, equipment, storage medium and product. Background Art

[0002] Driven by artificial intelligence, intelligent database search engines are constantly evolving. They have transcended the limitations of a single data format and are gradually acquiring multimodal data collection capabilities, enabling in-depth analysis of data semantics and dynamic anti-collection adaptability. These capabilities effectively address complex network environments and provide a solid technical foundation for efficient and accurate database search.

[0003] Faced with the need to develop customized search features, some users choose to develop their own data collection solutions. These solutions aim to leverage their own technical expertise to overcome challenges such as dynamic rendering, anti-collection countermeasures, and distributed scheduling, thereby enabling database data acquisition and providing a data foundation for search functionality, with the goal of creating a search service tailored to their business needs.

[0004] However, self-developed data collection solutions suffer from long development cycles and high maintenance costs. While relying on third-party services allows for regular data collection, the standardized collection logic conflicts significantly with the heterogeneous nature of local business data, requiring additional data processing and resulting in high communication costs. Summary of the Invention

[0005] The present application provides a data collection and scheduling method, apparatus, equipment, storage medium and product to solve the problem that when a third-party search engine service collects local database data, its standardized collection logic conflicts significantly with the heterogeneity of local business data, requiring additional data cleaning, format conversion and cross-system docking, resulting in a surge in communication costs.

[0006] In a first aspect, the present application provides a data collection and scheduling method, comprising:

[0007] Obtaining scheduling parameters for the task to be scheduled, and determining a corresponding configuration file based on the scheduling parameters;

[0008] generating a first data scheduling rule according to the configuration file;

[0009] Calling the data collection scheduling interface to execute the data collection scheduling task according to the first data scheduling rule;

[0010] When the data collection scheduling task is successfully executed, the scheduling data corresponding to the task to be scheduled is obtained.

[0011] In a second aspect, the present application provides a data acquisition and scheduling device, comprising:

[0012] An acquisition module is used to obtain the scheduling parameters of the task to be scheduled and determine the corresponding configuration file based on the scheduling parameters;

[0013] A generating module, configured to generate a first data scheduling rule according to the configuration file;

[0014] An execution module, configured to call a data collection scheduling interface and execute a data collection scheduling task according to the first data scheduling rule;

[0015] The determination module is used to obtain the scheduling data corresponding to the task to be scheduled when the data acquisition scheduling task is successfully executed.

[0016] In a third aspect, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;

[0017] The memory stores computer-executable instructions;

[0018] The processor executes the computer-executable instructions stored in the memory to implement the data acquisition scheduling method as described in the first aspect and various possible implementations of the first aspect.

[0019] In a fourth aspect, the present application provides a computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by a processor, are used to implement the data acquisition scheduling method as described in the first aspect and various possible implementations of the first aspect.

[0020] In a fifth aspect, the present application provides a program product, including a computer program, which implements the data acquisition scheduling method as described above when executed by a processor.

[0021] The data acquisition scheduling method, device, equipment, storage medium and product provided by this application, by constructing a dynamic parameter configuration engine, automatically generates adaptation rules from the configuration data, realizes the deployment of data acquisition strategies with one-click switching across environments, and effectively improves the adaptability and deployment efficiency of data acquisition strategies in different environments. Secondly, a unified interface abstraction layer is designed to encapsulate the scheduling interface required for interaction with third-party search engines as a standardized service, reducing the communication cost in the cross-team collaboration process. Finally, with the help of a unified scheduling interface, the adaptation rules are executed to complete the data acquisition scheduling task. In this way, the efficient transmission of data in multiple environments is guaranteed. At the same time, the time cost of technical personnel repeatedly adjusting data acquisition parameters and interface configurations due to differences in scheduling environments, database versions and security policies is reduced, thereby improving data acquisition efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0023] Figure 1 A flow chart of a data acquisition and scheduling method provided in an embodiment of the present application Figure 1 ;

[0024] Figure 2 A flow chart of a data acquisition and scheduling method provided in an embodiment of the present application Figure 2 ;

[0025] Figure 3 A flow chart of a data acquisition and scheduling method provided in an embodiment of the present application Figure 3 ;

[0026] Figure 4 A flow chart of a data acquisition and scheduling method provided in an embodiment of the present application Figure 4 ;

[0027] Figure 5 A schematic diagram of the structure of a data acquisition and scheduling device provided in this application;

[0028] Figure 6 This is a schematic diagram of the structure of an electronic device provided in this application.

[0029] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0030] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0031] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0032] In addition, this application involves conducting big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.), and using artificial intelligence technology to make automated decisions, and making technical solutions that have a significant impact on personal rights and interests based on the results of automated decisions. The application provides users with corresponding operation entrances for them to choose to agree or reject the results of automated decisions; if the user chooses to reject, the expert decision-making process will be entered.

[0033] It should be noted that the data collection and scheduling methods, devices, equipment, storage media and products provided in this application can be used in the field of big data, and can also be used in any field other than big data. The application fields of the data collection and scheduling methods, devices, equipment, storage media and products in this application are not limited.

[0034] Driven by artificial intelligence (AI), intelligent database search engines are gradually developing multimodal data collection, semantic parsing, and dynamic anti-collection adaptability. However, developing customized search capabilities requires significant resources to address dynamic rendering, anti-collection countermeasures, and distributed scheduling. This leads to long development cycles and high maintenance costs, presenting significant bottlenecks.

[0035] In view of this, more and more users choose to rely on scheduling third-party services for database searches. The search engine service provider collects local database data at regular intervals and uses the collected local database data agent to implement the search function. However, the standardized data collection logic conflicts significantly with the heterogeneity of local business data, requiring additional data cleaning, format conversion, and cross-system docking, resulting in a significant increase in communication costs. In addition, for technical personnel, due to IP isolation, database version differences, and inconsistent security policies in the development, testing, and production environments, data collection parameters and interface configurations need to be manually adjusted repeatedly. Cross-environment migration can easily cause data link breakage or performance degradation, seriously restricting R&D efficiency and system stability.

[0036] In response to the above problems, this application proposes a data collection and scheduling method, which automatically generates configuration files and adapted scheduling rules by building a dynamic parameter configuration engine, and encapsulates the scheduling interface required by third-party services into a standardized service. By calling the interface, the data collection scheduling task is executed according to the scheduling rules to obtain scheduling data, thereby realizing data collection scheduling in different scheduling environments under the same configuration file and the same interface, reducing labor costs, and avoiding data link breaks caused by scheduling environment migration.

[0037] This application can be applied to various business scenarios that require efficient data scheduling, such as scenarios where there are deployment requirements in multiple environments (such as development, testing, production, and other different operating environments).

[0038] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0039] Figure 1 A flow chart of a data acquisition and scheduling method provided in an embodiment of the present application Figure 1 .like Figure 1 As shown, the data collection scheduling method provided in this embodiment can be executed by a data collection system, including:

[0040] S101: Obtain scheduling parameters of a task to be scheduled, and determine a corresponding configuration file based on the scheduling parameters.

[0041] As you can understand, in a data acquisition system, the relevant scheduling parameters for the task to be scheduled are first obtained. These parameters define the configuration information required for task execution. Scheduling parameters typically include the Internet Protocol (IP) address segment corresponding to the environment in which the task will run (such as development, testing, or production), the type of database used, and the Application Programming Interface (API) version required for interaction with external systems. By parsing these parameters, the data acquisition system can determine and generate a configuration file suitable for the current task. This configuration file contains the detailed settings required for task execution, such as database connection information and API call addresses, thereby ensuring that the task can be executed correctly in different environments.

[0042] S102: Generate a first data scheduling rule according to the configuration file.

[0043] As you can understand, after determining the task configuration file, the data collection system can automatically generate the first data dispatch rule based on the information in the configuration file and combined with rule engine technology. A rule engine (such as an open source business rule engine) is a tool that automatically generates or adjusts business logic based on preset rules and conditions. Here, it dynamically generates data dispatch rules suitable for the current task based on parameters such as the database type and API version in the configuration file. These rules define how data is collected, processed, and stored, ensuring accurate and efficient data dispatch.

[0044] S103: calling a data collection scheduling interface and executing a data collection scheduling task according to a first data scheduling rule.

[0045] The data collection scheduling interface is the bridge between the data collection system and external data sources (such as third-party search engines). Through this interface, the data collection system can accurately transmit data collection requirements and instructions to the external data source and receive data returned from the external data source.

[0046] It is understandable that after generating the data scheduling rules, the data collection system can call the data collection scheduling interface to execute the data collection task according to the generated rules. The data scheduling rules include the specific requirements for data collection.

[0047] To simplify the calling process and improve system scalability, the data collection system leverages existing technologies, including synchronous service access components and declarative remote calling tools, to deeply abstract the scheduling interfaces of third-party service engines. The synchronous service access component simplifies the HTTP sending and response processing. The declarative remote calling tool defines interfaces through annotations, allowing developers to call remote services as if they were local methods.

[0048] By combining these two technologies, the data collection system abstracts the dispatch interface of third-party search engines into a unified Representational State Transfer (RESTful) service. RESTful services are a software architectural style based on the Hypertext Transfer Protocol (HTTP), which uses a unified interface and standard methods to operate resources. This allows the data collection system to call services through a single, unified interface, regardless of the underlying protocol (such as HTTP or HTTP Secure). This abstraction and unified interface design shields the differences in the underlying protocols, eliminating the need for extensive code modifications and adaptations when working with third-party search engines using different protocols, significantly improving system flexibility. Furthermore, this unified interface facilitates system maintenance, allowing technical staff to focus on implementing business logic rather than focusing on the details of the underlying protocols, thus improving system maintainability.

[0049] S104: When the data collection scheduling task is successfully executed, the scheduling data corresponding to the task to be scheduled is obtained.

[0050] As you can understand, when a data collection and scheduling task is successfully executed, the data collection system can obtain the scheduling data corresponding to the scheduled task. This data is the result of the task execution and may include web page content collected from search engines, records in databases, or other relevant information. After obtaining this data, the data collection system can further perform operations such as data processing, analysis, and storage to meet business needs.

[0051] This embodiment provides a data acquisition and scheduling method that obtains scheduling parameters for a task to be scheduled, determines a corresponding configuration file based on the scheduling parameters, and generates a first data scheduling rule based on the configuration file. The method then calls a data acquisition and scheduling interface to execute the data acquisition and scheduling task according to the first data scheduling rule. Upon successful execution of the data acquisition and scheduling task, the method obtains the scheduling data corresponding to the task to be scheduled. By constructing a dynamic parameter configuration engine to automatically generate configuration files based on the configuration parameters, the method enhances the automation and intelligence of data acquisition and scheduling, reduces manual intervention costs, and improves the efficiency and accuracy of data scheduling.

[0052] Figure 2 A flow chart of a data acquisition and scheduling method provided in an embodiment of the present application Figure 2 .like Figure 2 As shown, in Figure 1 Based on the embodiment, a method for generating a configuration file is described in detail, including:

[0053] S201: Perform a validity check on the scheduling parameters to obtain a first check result.

[0054] It can be understood that the scheduling parameters are the basis for generating the configuration file of the scheduling task. The scheduling parameters can be provided by the user. The correctness of the parameters is closely related to the execution of the subsequent scheduling tasks and data acquisition, so the scheduling parameters need to be verified for validity. The validity check can, for example, include verification of data types. For example, the basic types can be integers, floating-point numbers, strings, etc., to ensure that the type of the input parameters is consistent with expectations and to prevent program errors or security vulnerabilities caused by type mismatches; it can also include data range verification, which limits the value range of the parameters to ensure that they are within a reasonable range; it can also include data format verification, which verifies whether the parameters meet specific format requirements to ensure the standardization and parsability of the data; it can also include null and missing checks to ensure that the necessary parameters are not empty and that the required parameters have been provided. The specific verification method is not limited in this application.

[0055] S202: When the first verification result is verification passed, the scheduling parameters are configured into a preset configuration file to obtain a configuration file for the task to be scheduled.

[0056] As you can understand, if the scheduling parameters pass the initial legality check, then these parameters will be configured into the preset configuration file. The configuration file is the basis for the operation of the scheduled task, and it contains the core parameters and settings required for task execution. The preset configuration file can be formulated based on business needs, the performance of the target website for collection, the resources of the data collection system, legality and relevant regulations. In actual use, the specific collection task parameters can be filled in with the relevant data in the preset configuration file to obtain the corresponding configuration file. By correctly configuring the scheduling parameters into the preset configuration file, you can ensure that the task can be executed as expected.

[0057] S203: When the first verification result is verification failure, a parameter verification error log is recorded and the erroneous parameters in the scheduling parameters are corrected to obtain corrected parameters.

[0058] As can be understood, if the first verification result fails the validity check, the data acquisition system can record the relevant error information and correct the erroneous parameters, attempting to adjust them to values ​​that comply with the rules. Both the error information and the correction results can be recorded to form a parameter verification error log for technical personnel to access and review. The corrected parameters are the revised parameters and need to be verified again to ensure accuracy.

[0059] Among them, logs are recorded because a log system is integrated when configuring the data collection system. This log system can collect task execution logs in real time, mark the time consumption, data volume and abnormal information of each link, and the integrated log system includes a search engine, data processing and visualization platform.

[0060] S204: Perform a validity check on the correction parameters to obtain a second check result.

[0061] As you can understand, after the revised parameters are generated, they need to undergo another validation check to ensure that the revised parameters still comply with the pre-set rules. This step is similar to the initial validation check, but is specific to the revised parameters. The validation method is not detailed here. After the validation, a second validation result is obtained, indicating whether the revised parameters have passed validation.

[0062] S205: When the second verification result is verification passed, corresponding parameters in the scheduling parameters are replaced according to the correction parameters, and a configuration file is generated according to the replaced scheduling parameters.

[0063] It is understandable that when the revised parameters pass the legality check, it means that these revised parameters meet the requirements of the data acquisition system. The data acquisition system can accurately locate the parameters in the original scheduling parameters that were judged to be incorrect during the first check, and replace them one by one with the revised parameters that have now passed the check. This replacement process needs to ensure the accuracy of the parameter mapping to avoid errors in the parameter correspondence. After the replacement is completed, the data acquisition system can create a configuration file for the task to be scheduled based on the updated scheduling parameter set containing the correct parameters, according to the preset configuration file generation rules and format.

[0064] S206: When the second verification result is verification failure, the current number of corrections and the parameter verification error log are recorded.

[0065] As you can understand, if a modified parameter fails the validity check, the data acquisition system can immediately record the number of corrections that have been made so far and also record in detail all error messages that occurred during the parameter verification process, including but not limited to which parameters did not comply with the rules and the specific manifestations of the non-compliance, thus forming a complete parameter verification error log. This log information is crucial for subsequent troubleshooting and analysis.

[0066] S207: Determine whether the current number of corrections is equal to the preset number. If so, execute step S209; if not, execute step S208.

[0067] As you can understand, after recording the relevant information, the data collection system can determine whether the current number of correction attempts has reached the preset maximum number. The preset number is determined based on a combination of factors such as actual business needs, system performance, and the complexity of parameter correction. It is intended to avoid wasting system resources and degrading performance due to unlimited correction attempts.

[0068] S208: Correct the erroneous parameters in the correction parameters again to obtain new correction parameters, and return to execute the verification process for the correction parameters.

[0069] As you can understand, if the number of corrections does not reach the preset number and the correction parameter verification fails, the data acquisition system can initiate a new round of correction operations. First, it conducts an in-depth analysis of the erroneous parameters still present in the correction parameters. In combination with system rules and business logic, a more precise correction strategy is used to correct these erroneous parameters again, hoping to obtain new correction parameters that comply with the rules. Subsequently, the validity verification of the correction parameters can be performed again, forming a closed-loop correction and verification process until the parameters pass verification or the maximum number of corrections is reached.

[0070] S209: Output all parameter verification error logs.

[0071] Understandably, when the number of corrections reaches the preset maximum number but the corrected parameters still fail the legality check, the data acquisition system can assume that sufficient correction attempts have been made according to the established rules, but the parameters still cannot meet the requirements. At this time, the data acquisition system can comprehensively organize and output all parameter verification error logs generated from the first verification to all current correction attempts. These logs contain key information such as the details of the parameters that failed each verification, the error type, the correction attempts, and the final results. This provides a detailed and reliable basis for subsequent troubleshooting, parameter adjustment, or system optimization, helping to fundamentally solve the problem and improve the stability and reliability of the scheduling system.

[0072] This embodiment provides a data collection and scheduling method that begins with a preliminary validity check, obtains a first verification result, generates a configuration file or modifies configuration parameters based on the first verification result, re-verifies the modified parameters, and finally generates a corresponding configuration file or outputs an error log based on a second verification result. This entire process ensures the validity and correctness of the parameters and can automatically modify the configuration parameters within a limited number of times. This method can reduce the occurrence of task execution failures or abnormal situations caused by parameter errors, and the detailed error log also provides strong support for subsequent problem investigation and repair.

[0073] Figure 3 A flow chart of a data acquisition and scheduling method provided in an embodiment of the present application Figure 3 .like Figure 3 As shown, in Figure 1 Based on the embodiment, a method for adjusting the execution error of the data acquisition scheduling task is described in detail, including:

[0074] S301: In the event of an execution error in a data collection scheduling task, the execution log of the target historical task is retrieved.

[0075] The target historical task is a historical task in the same scheduling environment as the current data collection scheduling task.

[0076] Understandably, when errors occur during the execution of a data collection and scheduling task, the data collection system can retrieve the execution logs of historical tasks that were run under the same scheduling environment as the current task in order to identify the root cause and develop a suitable solution. The same scheduling environment means that these historical tasks ran under the same hardware resources, network conditions, and software configuration as the current task. Their execution logs are highly valuable for analyzing the current task's problems.

[0077] Combined with the historical data analysis model, the data collection system can use the streaming data processing framework to perform aggregate analysis on the acquired historical task execution logs. This framework can process large-scale data streams in real time, has high throughput and low latency, and can efficiently complete the aggregation calculation tasks of log data. By analyzing the execution of historical tasks in different time periods, such as the difference in execution efficiency during peak and off-peak periods, the load characteristics of the tasks can be dynamically understood. Based on this, data support can be provided for subsequent analysis of current task anomalies. For example, based on the load conditions during historical peak periods, it can be preliminarily determined whether the errors of the current task during peak periods are related to excessive load, thus laying the foundation for further determining the cause of the anomaly.

[0078] S302: Determine the abnormal log of the current data collection scheduling task.

[0079] As you can understand, the exception log for the current data collection and scheduling task records various exception information that occurs during task execution, such as error codes and exception stack traces. This information directly reflects the problems encountered during task execution and their location. By analyzing the exception log, we can initially locate the approximate scope of the problem. Using the streaming data processing framework to monitor and analyze the log data of the current task in real time, when an abnormal log is detected, we can quickly mark and extract key exception information.

[0080] At the same time, combining historical data analysis with comparing current anomaly logs with similar anomalies in previous tasks can help more accurately determine the nature and severity of the anomaly. For example, if a previous task experienced a similar anomaly under the same circumstances and the anomaly was resolved by adjusting the collection frequency, then the current task may also need to consider analyzing the collection frequency.

[0081] S303: Based on the exception log, analyze and process the execution log to obtain the exception cause and determine the adjustment strategy.

[0082] Understandably, after identifying the current task's exception log, further in-depth analysis is necessary in conjunction with the execution logs of historical tasks. Execution logs record the entire execution process of historical tasks from start to finish, including status at each stage and resource usage. By comprehensively analyzing both exception and execution logs, the root cause of the current task's execution error can be identified.

[0083] For example, a machine learning model based on a multilayer perceptron (MLP) can be used to analyze historical task performance data, as well as exception and execution logs for current tasks. This model can learn the relationship between different exception causes and various parameters, such as the correlation between parameters like thread pool size and timeout period and task execution success rate and efficiency. This model analysis enables more scientific determination of adjustment strategies. For example, based on historical data, the model can predict to what extent adjusting the thread pool size will effectively improve task execution efficiency under the current exception situation, providing a more accurate basis for determining adjustment strategies.

[0084] S304: Adjust the resource allocation parameters of the current data acquisition scheduling task according to the adjustment strategy to obtain new resource allocation parameters.

[0085] As you can understand, after determining the adjustment strategy, you can adjust the resource allocation parameters of the current data acquisition scheduling task accordingly. Resource allocation parameters can include thread pool size, memory allocation, CPU usage limit, etc. The proper setting of these parameters is crucial for the smooth execution of the task. Adjusting parameters according to the adjustment strategy ensures that the task runs in an appropriate resource environment, avoiding task execution errors or inefficiencies caused by insufficient or wasted resources.

[0086] During implementation, the dynamic parameter tuning mechanism in the data acquisition system can be leveraged to analyze historical task performance using an MLP-based machine learning model to automatically optimize parameters such as thread pool size and timeouts. Based on the performance of historical tasks under different resource allocation scenarios, the model can predict the performance of the current task under adjusted resource allocation parameters. For example, the model might recommend increasing the thread pool size from 10 to 15 threads to address the concurrent processing pressure of the current task, thereby obtaining new, more reasonable resource allocation parameters.

[0087] S305: Generate a second data scheduling rule according to the new resource allocation parameter and configuration file, and re-execute the data collection scheduling task according to the second data scheduling rule.

[0088] As you can understand, after obtaining the new resource allocation parameters, you can combine the original configuration file to generate a new data scheduling rule, namely the second data scheduling rule. The new scheduling rule will comprehensively consider factors such as resource allocation, task priority, and data source characteristics to ensure that the scheduled task can be re-executed according to more reasonable rules.

[0089] It should be noted that if the data collection scheduling task has errors multiple times at the same location and has not reached the preset maximum number of corrections, the data collection system can correct it multiple times. If the data collection scheduling task still has errors at the same location after reaching the preset maximum number of corrections, all current execution logs of the data collection scheduling task will be output so that technical personnel can analyze the cause of the error in the scheduling task based on the log information and make targeted corrections.

[0090] This embodiment provides a data acquisition scheduling method. When a data acquisition scheduling task fails, it compares the execution logs of the target historical tasks under the same scheduling environment with the exception logs of the current task to identify the cause of the exception and determine an adjustment strategy. The method then adjusts the resource allocation parameters of the current task to obtain new parameters. This method can quickly locate the cause of the task execution error and effectively resolve the exception by dynamically adjusting the resource allocation parameters. This improves the success rate and stability of data acquisition scheduling tasks, reduces resource waste and time delays caused by task failures, and enhances the fault tolerance and self-repair capabilities of the data acquisition system.

[0091] Figure 4 A flow chart of a data acquisition and scheduling method provided in an embodiment of the present application Figure 4 .like Figure 4 As shown, in Figure 1 Based on the embodiment, the execution process of one-key switching scheduling environment is described in detail, including:

[0092] S401: Receive a user's scheduling environment change request.

[0093] It's understandable that the data acquisition system can be integrated into a visualization device. Through the monitoring panel provided by the visualization device, users can enter configuration data and change the scheduling environment. For example, a data acquisition scheduling task originally running in a development environment can be switched to a test environment. During the data acquisition system integration process, for example, a monitoring panel can be constructed using the monitoring system and visualization platform. This panel can display task progress, resource utilization, and data transfer success rate in real time, helping users understand the system's operating status.

[0094] In the data acquisition system, users can adjust the execution environment of tasks based on their needs. The data acquisition system can receive scheduling environment change requests through the monitoring panel. These requests may include new environment identifiers, specific configuration parameters, or other relevant information. Upon receiving these requests, the data acquisition system can perform preliminary verification and processing to ensure the legitimacy and validity of the requests.

[0095] S402: Modify the configuration file according to the scheduling environment change request, and generate a third data scheduling rule according to the modified configuration file.

[0096] As you can understand, the data collection system needs to execute data collection scheduling tasks based on the configuration file. After receiving a user's scheduling change request, the data collection system can modify the corresponding data in the configuration file based on the information in the request. The modifications may include: environmental parameters, collection strategy adjustments, and logging. After the configuration file is modified, the data retrieval rules can be regenerated based on the new configuration information to determine the specific execution method of the task.

[0097] S403: Execute a new data collection scheduling task according to the third data scheduling rule through the data collection scheduling interface.

[0098] It can be understood that, finally, the data acquisition system can apply the generated new data scheduling rules to the actual data acquisition scheduling tasks according to the unified scheduling interface.

[0099] This embodiment provides a data acquisition and scheduling method, which receives a user's scheduling environment change request, modifies a configuration file according to the scheduling environment change request, and generates a third data scheduling rule according to the modified configuration file. Through the data acquisition and scheduling interface, a new data acquisition scheduling task is executed according to the third data scheduling rule, avoiding manual modification of the scheduling code or scheduling configuration file, providing highly reliable, full-link automated data acquisition support for the business, and realizing one-click conversion of the scheduling environment, solving the problem of data link breakage or performance degradation that is easily caused during cross-environment migration.

[0100] Figure 5 This is a structural diagram of a data acquisition and scheduling device provided by this application. Figure 5 As shown, the present application provides a data acquisition and scheduling device, and the data acquisition and scheduling device 500 includes:

[0101] The acquisition module 501 is used to obtain the scheduling parameters of the task to be scheduled and determine the corresponding configuration file based on the scheduling parameters;

[0102] A generating module 502 is configured to generate a first data scheduling rule according to the configuration file;

[0103] An execution module 503 is configured to call a data collection scheduling interface and execute a data collection scheduling task according to the first data scheduling rule;

[0104] The determination module 504 is configured to obtain the scheduling data corresponding to the task to be scheduled if the data acquisition scheduling task is successfully executed.

[0105] Optionally, the determination module 504 is specifically used to perform a validity check on the scheduling parameters to obtain a first verification result; if the first verification result is a passed verification, the scheduling parameters are configured into a preset configuration file to obtain the configuration file of the task to be scheduled.

[0106] Optionally, the device further includes: a processing module 505;

[0107] The processing module 505 is used to, when the first verification result is verification failure, record a parameter verification error log and correct the erroneous parameters in the scheduling parameters to obtain corrected parameters; perform a validity verification on the corrected parameters to obtain a second verification result; when the second verification result is verification passing, replace the corresponding parameters in the scheduling parameters according to the corrected parameters, and generate the configuration file based on the replaced scheduling parameters.

[0108] Optionally, the device further includes: a judgment module 506;

[0109] The judgment module 506 is further configured to, when the second verification result is a verification failure, record the current number of corrections and a parameter verification error log, and determine whether the current number of corrections is equal to a preset number;

[0110] The processing module 505 is further configured to correct the erroneous parameters in the correction parameters again to obtain new correction parameters if the current correction number is not equal to the preset number, and return to execute the verification process for the correction parameters;

[0111] The generating module 502 is further configured to output all parameter verification error logs if the current number of corrections is equal to the preset number of corrections.

[0112] Optionally, the processing module 505 is further configured to optimize the resource allocation parameters of the current data acquisition scheduling task according to the historical scheduling tasks to obtain new resource allocation parameters when an error occurs in the execution of the data acquisition scheduling task.

[0113] The generating module 502 is further configured to generate a second data scheduling rule according to the new resource allocation parameter and the configuration file, and re-execute the data collection scheduling task according to the second data scheduling rule.

[0114] Optionally, the processing module 505 is specifically used to retrieve the execution log of the target historical task, where the target historical task is a historical task in the same scheduling environment as the current data acquisition scheduling task; determine the exception log of the current data acquisition scheduling task; based on the exception log, analyze and process the execution log to obtain the cause of the exception and determine the adjustment strategy; adjust the resource allocation parameters of the current data acquisition scheduling task according to the adjustment strategy to obtain new resource allocation parameters.

[0115] Optionally, the acquisition module 501 is further configured to receive a scheduling environment change request from a user;

[0116] The generating module 502 is further configured to modify the configuration file according to the scheduling environment change request, and generate a third data scheduling rule according to the modified configuration file;

[0117] The execution module 503 is further configured to execute a new data collection scheduling task according to the third data scheduling rule through the data collection scheduling interface.

[0118] The data acquisition and scheduling device provided in the embodiment of the present application has an implementation principle and technical effects similar to the implementation methods of the various parts of the aforementioned data acquisition and scheduling method, and will not be repeated here.

[0119] Figure 6 This is a schematic diagram of the structure of an electronic device provided by this application. Figure 6 As shown, the present application provides an electronic device, which includes a receiver 601, a transmitter 602, a processor 603 and a memory 604.

[0120] Receiver 601, for receiving instructions and data;

[0121] Transmitter 602, used to send instructions and data;

[0122] Memory 604, for storing computer-executable instructions;

[0123] The processor 603 is configured to execute the computer-executable instructions stored in the memory 604 to implement the various steps of the data acquisition and scheduling method in the above embodiment. For details, please refer to the relevant description in the above embodiment of the data acquisition and scheduling method.

[0124] Optionally, the memory 604 may be independent or integrated with the processor 603 .

[0125] When the memory 604 is independently provided, the electronic device further includes a bus for connecting the memory 604 and the processor 603 .

[0126] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the aforementioned embodiments and will not be described in detail here.

[0127] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method described in any of the above embodiments is implemented.

[0128] An embodiment of the present application further provides a computer program product, including a computer program, which implements the method described in any of the aforementioned embodiments when executed by a processor.

[0129] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.

[0130] It should be further noted that, although the various steps in the flowchart are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps may be performed in other orders. Moreover, at least a portion of the steps in the flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but may be performed at different times. The execution order of these sub-steps or stages is not necessarily to be performed in sequence, but may be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0131] It should be understood that the above-described device embodiments are merely illustrative, and the device of the present application may also be implemented in other ways. For example, the division of units / modules in the above-described embodiments is merely a logical functional division, and actual implementations may employ other division methods. For example, multiple units, modules, or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0132] In addition, unless otherwise specified, the functional units / modules in the various embodiments of the present application may be integrated into a single unit / module, each unit / module may exist physically separately, or two or more units / modules may be integrated together. The aforementioned integrated units / modules may be implemented in the form of hardware or software program modules.

[0133] If an integrated unit / module is implemented in hardware, the hardware may be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor may be any appropriate hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC. Unless otherwise specified, the storage unit may be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.

[0134] If the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes: U disk, read-only memory (ROM), random access memory (RAM), mobile hard disk, magnetic disk, or optical disk, etc., various media that can store program code.

[0135] In the above embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined in any way. To keep the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0136] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0137] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A data collection and scheduling method, characterized in that: include: Obtaining scheduling parameters for the task to be scheduled, and determining a corresponding configuration file based on the scheduling parameters; generating a first data scheduling rule according to the configuration file; Calling the data collection scheduling interface to execute the data collection scheduling task according to the first data scheduling rule; When the data collection scheduling task is successfully executed, the scheduling data corresponding to the task to be scheduled is obtained.

2. The method according to claim 1, characterized in that The determining a corresponding configuration file based on the scheduling parameter includes: Performing a validity check on the scheduling parameters to obtain a first check result; If the first verification result is that the verification is passed, the scheduling parameters are configured into a preset configuration file to obtain the configuration file of the task to be scheduled.

3. The method according to claim 2, characterized in that The method further comprises: If the first verification result is a verification failure, recording a parameter verification error log and correcting the erroneous parameters in the scheduling parameters to obtain corrected parameters; Performing a validity check on the correction parameter to obtain a second check result; When the second verification result is that the verification is passed, the corresponding parameters in the scheduling parameters are replaced according to the correction parameters, and the configuration file is generated according to the replaced scheduling parameters.

4. The method according to claim 3, characterized in that The method further comprises: If the second verification result is a verification failure, record the current number of corrections and the parameter verification error log, and determine whether the current number of corrections is equal to the preset number; If the current number of corrections is not equal to the preset number, the erroneous parameters in the correction parameters are corrected again to obtain new correction parameters, and the process of verifying the correction parameters is returned to be executed; If the current number of corrections is equal to the preset number, all parameter verification error logs are output.

5. The method according to claim 1, wherein The method further comprises: In the event of an execution error of the data acquisition scheduling task, optimizing the resource allocation parameters of the current data acquisition scheduling task according to the historical scheduling tasks to obtain new resource allocation parameters; A second data scheduling rule is generated according to the new resource allocation parameter and the configuration file, and the data collection scheduling task is re-executed according to the second data scheduling rule.

6. The method according to claim 5, characterized in that The resource allocation parameters of the current data acquisition scheduling task are optimized based on the historical scheduling tasks to obtain new resource allocation parameters, including: Retrieve the execution log of the target historical task, where the target historical task is a historical task in the same scheduling environment as the current data collection scheduling task; Determine the abnormal log of the current data acquisition scheduling task; Based on the exception log, the execution log is analyzed and processed to obtain the cause of the exception and determine the adjustment strategy; The resource allocation parameters of the current data acquisition scheduling task are adjusted according to the adjustment strategy to obtain new resource allocation parameters.

7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: Receive scheduling environment change requests from users; Modifying the configuration file according to the scheduling environment change request, and generating a third data scheduling rule according to the modified configuration file; A new data collection scheduling task is executed according to the third data scheduling rule through the data collection scheduling interface.

8. A data acquisition and scheduling device, characterized in that: include: An acquisition module is used to obtain the scheduling parameters of the task to be scheduled and determine the corresponding configuration file based on the scheduling parameters; A generating module, configured to generate a first data scheduling rule according to the configuration file; An execution module, configured to call a data collection scheduling interface and execute a data collection scheduling task according to the first data scheduling rule; The determination module is used to obtain the scheduling data corresponding to the task to be scheduled when the data acquisition scheduling task is successfully executed.

9. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.

11. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 7 when being executed by a processor.