An intelligent monitoring system and method for task logs of offline data development based on large models
By building an intelligent monitoring system for offline data development task logs based on large models, the problem of dirty data affecting synchronous data capabilities is solved, and the automated monitoring and management of dirty data is realized, which improves the success rate and efficiency of offline data synchronization.
Patent Information
- Application Number
- CN202411069156.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-08-06
AI Technical Summary
During the offline data synchronization process of software development, the generation of dirty data causes tasks to fail or be discarded, affecting the ability and stability of synchronized data. The existing technology cannot effectively solve the automated monitoring and management of dirty data.
By building an intelligent monitoring system for offline data development task logs based on large models, including data integration directory module, offline synchronization module, production environment operation and maintenance management module and intelligent monitoring module, automated monitoring and management of dirty data are realized, dirty data processing nodes are built based on sample data, and the start and stop of synchronization tasks are controlled.
The data synchronization of offline tasks is realized, the risk of error reporting is reduced, the write success rate and efficiency is improved, and the digital synchronization level of software development is improved.
Smart Images

Figure CN119025375B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent monitoring of software development, and particularly to an intelligent monitoring system and method for offline data development task logs based on a large model. Background Technique
[0002] During the offline data synchronization process of software development, the data of the data source at the source end needs to be written into the data source at the destination end, that is, the data type at the source end needs to match the data type at the write end. However, during the offline task process, a large amount of dirty data is usually generated. That is, if an exception occurs during the process of writing a single piece of data into the target data source, then this piece of data is dirty data. Generally, the data that fails to be written is classified as dirty data. In the actual operation process, the administrator can set the quantity threshold for the generation of dirty data through the task to ensure that the synchronization task configuration in the data integration process completes the data synchronization. When the number of generated dirty data exceeds the quantity threshold, the task will fail and exit. For example, if the allowable number of dirty data is set to 0, then when dirty data is generated, the task will fail and exit. When the number of generated dirty data is less than the quantity threshold, the task will continue to run, but the dirty data will be discarded and not written to the destination end.
[0003] The setting of the task number of dirty data is generally manually set by the administrator. However, during the learning process of the large model, due to the different natures of the offline data, using a stable task number will also affect the ability to synchronize data and various types of error reports will occur. Summary of the Invention
[0004] The purpose of the present invention is to provide an intelligent monitoring system and method for offline data development task logs based on a large model to solve the problems raised in the above background technique.
[0005] To solve the above technical problems, the present invention provides the following technical solution: An intelligent monitoring method for offline data development task logs based on a large model, the method includes the following steps:
[0006] S1. Obtain the name of the software development business process, drag the offline synchronization node in the data integration directory module to construct an offline synchronization task;
[0007] S2. Under the offline synchronization task, configure the synchronization network link, select the data source and data destination of the offline synchronization task, as well as the resource group for executing the synchronization task, test the connectivity, edit the script, and configure the offline synchronization task;
[0008] S3. Import the offline synchronization task into the production environment operation and maintenance center, and the production environment operation and maintenance center calls the large model to configure the dirty data processing node;
[0009] S4. According to the dirty data processing node and the offline synchronization task, perform intelligent monitoring of the offline data development task log, and control the start and stop of the synchronization task of the offline data development task log.
[0010] According to the above technical solution, in step S2, it further includes:
[0011] Select the switch network segment bound by the resource group, the EIP of the old version resource group itself, or the EIP configured by the VPC bound by the new version general-purpose resource group, and add it to the whitelist of data sources;
[0012] Build the data association between the data source and DataWorks, and add the DataWorks data source to test connectivity. Specifically, it includes:
[0013] Under the data integration directory module, log in to the DataWorks console, select data integration in the navigation bar, enter data integration after selecting the corresponding workspace, select the data source in the navigation bar, then add a new data source in the data source list, set the new data source as the current data source, and configure relevant connection parameters to achieve testing connectivity;
[0014] The configuration of the offline synchronization task further includes:
[0015] The offline synchronization task with periodic scheduling needs to configure relevant attributes when the task is automatically scheduled. The relevant attributes include: configuring node scheduling attributes, configuring time attributes, and configuring resource attributes;
[0016] The configuration of the node scheduling attributes includes: assigning scheduling parameters to custom variables, and at the same time supporting the assignment of constants;
[0017] The configuration of the time attributes includes: being used to define the periodic scheduling method of the task in the production environment;
[0018] The configuration of the resource attributes includes: in the scenario of the offline synchronization task with periodic scheduling, defining the scheduling resource group used when executing the data integration task resources.
[0019] According to the above technical solution, in step S3, it further includes:
[0020] Based on different resource groups for executing the synchronization task, build large models under different resource groups respectively. The method for building a large model under any resource group includes:
[0021] Construct a sample data set, select the original data under the same resource group, where the original data refers to the manual setting data of the user administrator, and form a specific data set [m0, y0]. Among them, m0 refers to the offline synchronization task feature group, specifically including the total number of offline synchronization tasks and the number of associations between offline synchronization tasks. The association means that there is data interaction between any two offline synchronization tasks, which is defined as the number of associations; y0 refers to the number of configured dirty data processing nodes; Assign values to m0 in a fixed weight format, including:
[0022] m0 = z1 * a0 + (1 - z1) * b0
[0023] Among them, z1 refers to the weight set by the system; a0 and b0 respectively represent the normalized data of the total number of offline synchronization tasks and the number of associations between offline synchronization tasks;
[0024] Based on the sample data set, construct a frequency histogram. In the frequency histogram, the area of the rectangle is defined as the interval frequency, and the height of the rectangle is the average frequency density of the interval. Then:
[0025]
[0026] Among them, x i refers to the sample data i in the sample data set; f refers to the probability density function based on the sample data set; F refers to the cumulative distribution function based on the sample data set; h represents the bandwidth of the kernel density estimation, set by the system; F(x i +h), F(x i -h) respectively represent the cumulative distribution function values under (x i +h), (x i -h), that is, for a continuous function, the sum of the probabilities of all values less than or equal to (x i +h) is the cumulative distribution function value under (x i +h); the sum of the probabilities of all values less than or equal to (x i -h) is the cumulative distribution function value under (x i -h);
[0027] In the two-dimensional scenario of the specific data set [m0, y0], for any new data x, find its closest range [x i -h, x i +h] in the sample data set; form the calculation of the probability density function for any new data x:
[0028]
[0029] Among them, s represents the serial number; n represents the total number of samples in the sample data set; K0 represents the uniform kernel function; dis(x, [m s 、ys ) represents the distance between x and [m s , y s .
[0030] Based on the total number of offline synchronization total tasks and the association quantity between offline synchronization tasks monitored by the system, a specific new data set is formed, thereby outputting the number of configured dirty data processing nodes under the specific new data set.
[0031] According to the above technical solution, it further includes:
[0032] Obtain the offline synchronization task and the number of configured dirty data processing nodes under the offline synchronization task, and perform intelligent monitoring of the offline data development task log. If the number of dirty data generated within the period under the configured time attribute is less than or equal to the number of configured dirty data processing nodes, the synchronization task of the offline data development task log is normally carried out; if the number of dirty data generated within the period under the configured time attribute is higher than the number of configured dirty data processing nodes, the synchronization task of the offline data development task log is stopped, the synchronized offline task data is retained, and an error is reported at the administrator port simultaneously.
[0033] An intelligent monitoring system for offline data development task logs based on a large model, the system includes: a data integration directory module, an offline synchronization module, a production environment operation and maintenance management module, and an intelligent monitoring module;
[0034] The data integration directory module is used to obtain the software development business process name, drag the offline synchronization node in the data integration directory module to construct an offline synchronization task; the offline synchronization module is used to configure the synchronization network link under the offline synchronization task, select the data source and data destination of the offline synchronization task, and the resource group for executing the synchronization task, test the connectivity, edit the script, and configure the offline synchronization task; the production environment operation and maintenance management module is used to import the offline synchronization task into the production environment operation and maintenance center, and the production environment operation and maintenance center calls the large model to configure the dirty data processing node; the intelligent monitoring module is used to perform intelligent monitoring of the offline data development task log according to the dirty data processing node and the offline synchronization task, and control the start and stop of the synchronization task of the offline data development task log;
[0035] The output end of the data integration directory module is connected to the input end of the offline synchronization module; the output end of the offline synchronization module is connected to the input end of the production environment operation and maintenance management module; the output end of the production environment operation and maintenance management module is connected to the input end of the intelligent monitoring module.
[0036] According to the above technical solution, the data integration directory module includes a software development business unit and an offline synchronization task construction unit;
[0037] The software development business unit is used to record the names of software development business processes and form a software development business list; the offline synchronization task construction unit is used to drag offline synchronization nodes to construct an offline synchronization task;
[0038] The output end of the software development business unit is connected to the input end of the offline synchronization task construction unit.
[0039] According to the above technical solution, the offline synchronization module includes an offline synchronization test unit and an offline synchronization configuration unit;
[0040] The offline synchronization test unit is used to configure a synchronization network link under an offline synchronization task, select the data source and data destination of the offline synchronization task, and the resource group for executing the synchronization task, and test the connectivity; the offline synchronization configuration unit is used to edit a script and configure relevant attributes during the automatic scheduling of a task under a periodically scheduled offline synchronization task, and the relevant attributes include: configuring node scheduling attributes, configuring time attributes, and configuring resource attributes;
[0041] The output end of the offline synchronization test unit is connected to the input end of the offline synchronization configuration unit.
[0042] According to the above technical solution, the production environment operation and maintenance management module includes a production environment operation and maintenance center and a large model analysis unit;
[0043] The production environment operation and maintenance center is used to import an offline synchronization task and call the large model for processing; the large model analysis unit forms a large model under different resource groups based on sample data;
[0044] The output end of the large model analysis unit is connected to the input end of the production environment operation and maintenance center.
[0045] According to the above technical solution, the intelligent monitoring module includes an intelligent monitoring unit and an automatic start-stop control unit;
[0046] The intelligent monitoring unit is used to perform intelligent monitoring of the offline data development task log according to the dirty data processing node and the offline synchronization task, and control the start and stop of the synchronization task of the offline data development task log; the automatic start-stop control unit is used to obtain the offline synchronization task and the number of configured dirty data processing nodes under the offline synchronization task, perform intelligent monitoring of the offline data development task log, if the number of dirty data generated within the period under the configured time attribute is less than or equal to the number of configured dirty data processing nodes, the synchronization task of the offline data development task log is carried out normally; if the number of dirty data generated within the period under the configured time attribute is higher than the number of configured dirty data processing nodes, the synchronization task of the offline data development task log is stopped, the already synchronized offline task data is retained, and an error is reported at the administrator port at the same time.
[0047] Compared with the prior art, the beneficial effects achieved by the present invention are as follows: The present invention can realize data synchronization of offline tasks during the software development process, and construct the quantity nodes of dirty data based on the large model under sample data to form automated monitoring and management. During the synchronous writing process of the offline data development task log, the error reporting risk is reduced, the writing success rate and efficiency are improved, the digital offline synchronization level is achieved, and the software development ability is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention.
[0049] In the drawings:
[0050] Figure 1 is a schematic flow chart of an intelligent monitoring method for offline data development task logs based on a large model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0052] Please refer to Figure 1 , in the first embodiment, an intelligent monitoring method for offline data development task logs based on a large model is provided. The method includes: obtaining the name of the software development business process, dragging the offline synchronization node in the data integration directory module to construct an offline synchronization task; under the offline synchronization task, configuring the synchronization network link, selecting the data source and data destination of the offline synchronization task, as well as the resource group for executing the synchronization task, testing the connectivity, editing the script, and configuring the offline synchronization task;
[0053] Select the switch network segment bound by the resource group, the EIP of the old version resource group itself, or the EIP configured by the VPC bound by the new version general-purpose resource group, and add it to the whitelist of the data source;
[0054] Construct the data association between the data source and DataWorks, and test the connectivity by adding the DataWorks data source. Specifically, it includes:
[0055] Under the data integration catalog module, log in to the DataWorks console. In the navigation bar, select Data Integration. After selecting the corresponding workspace, enter Data Integration. In the navigation bar, select Data Sources, and then add a new data source in the data source list. Set the newly added data source as the current data source and configure the relevant connection parameters to achieve test connectivity;
[0056] The configuration of the offline synchronization task further includes:
[0057] For the offline synchronization task with periodic scheduling, relevant attributes for automatic task scheduling need to be configured. The relevant attributes include: configuring node scheduling attributes, configuring time attributes, and configuring resource attributes;
[0058] The configuration of node scheduling attributes includes: assigning scheduling parameters to custom variables, and at the same time supporting the assignment of constants;
[0059] The configuration of time attributes includes: defining the periodic scheduling method of the task in the production environment;
[0060] The configuration of resource attributes includes: in the scenario of the offline synchronization task with periodic scheduling, defining the scheduling resource group used when defining the execution resources of the data integration task.
[0061] Import the offline synchronization task into the production environment operation and maintenance center, and the production environment operation and maintenance center calls the large model to configure the dirty data processing node;
[0062] Based on different resource groups for executing synchronization tasks, large models under different resource groups are respectively constructed. The method for constructing a large model under any resource group includes:
[0063] Construct a sample data set, select the original data under the same resource group. The original data refers to the manually set data by the user administrator, forming a specific data set [m0, y0], where m0 refers to the offline synchronization task feature group, specifically including the total number of offline synchronization tasks and the number of associations between offline synchronization tasks. The association means that there is data interaction between any two offline synchronization tasks, which is defined as one association number; y0 refers to the number of configured dirty data processing nodes; Assign values to m0 in a fixed weight format, including:
[0064] m0 = z1 * a0+(1 - z1) * b0
[0065] Among them, z1 refers to the weight set by the system; a0 and b0 respectively represent the normalized data of the total number of offline synchronization tasks and the number of associations between offline synchronization tasks;
[0066] Based on the sample data set, construct a frequency histogram. In the frequency histogram, the area of the rectangle is defined as the interval frequency, and the height of the rectangle is the average frequency density of the interval. Then:
[0067]
[0068] Among them, x i refers to the sample data i in the sample data set; f refers to the probability density function based on the sample data set; F refers to the cumulative distribution function based on the sample data set; h represents the bandwidth of kernel density estimation, which is set by the system; F(x i +h), F(x i -h) respectively represent the cumulative distribution function values at (x i +h), (x i -h), that is, for a continuous function, the sum of the occurrence probabilities of all values less than or equal to (x i +h) is the cumulative distribution function value at (x i +h); the sum of the occurrence probabilities of all values less than or equal to (x i -h) is the cumulative distribution function value at (x i -h);
[0069] In the two-dimensional scenario of a specific data set [m0, y0], for any new data x, find its closest range [x i -h, x i +h] in the sample data set; form the calculation of the probability density function for any new data x:
[0070]
[0071] Among them, s represents the serial number; n represents the total number of samples in the sample data set; K0 represents the uniform kernel function; dis(x, [m s , y s ) represents the distance between x and [m s , y s ;
[0072] According to the total number of offline synchronization total tasks and the association quantity between offline synchronization tasks monitored by the system, form a specific new data set, and thus output the number of configured dirty data processing nodes under the specific new data set.
[0073] It also includes:
[0074] Obtain the offline synchronization tasks and the number of configured dirty data processing nodes under the offline synchronization tasks, and conduct intelligent monitoring of the offline data development task logs. If the number of dirty data generated within the period under the configured time attribute is less than or equal to the number of configured dirty data processing nodes, the synchronization task of the offline data development task logs is normally carried out; if the number of dirty data generated within the period under the configured time attribute is higher than the number of configured dirty data processing nodes, the synchronization task of the offline data development task logs is stopped, the already synchronized offline task data is retained, and an error is reported at the administrator port simultaneously.
[0075] In the second embodiment, an intelligent monitoring system for the task log of offline data development based on a large model is provided. The system includes: a data integration directory module, an offline synchronization module, a production environment operation and maintenance management module, and an intelligent monitoring module;
[0076] The data integration directory module is used to obtain the software development business process name, drag and drop the offline synchronization node in the data integration directory module to construct an offline synchronization task; the offline synchronization module is used to configure the synchronization network link under the offline synchronization task, select the data source and data destination of the offline synchronization task, and the resource group for executing the synchronization task, test the connectivity, edit the script, and configure the offline synchronization task; the production environment operation and maintenance management module is used to import the offline synchronization task into the production environment operation and maintenance center, and the production environment operation and maintenance center calls the large model to configure the dirty data processing node; the intelligent monitoring module is used to perform intelligent monitoring of the task log of the offline data development according to the dirty data processing node and the offline synchronization task, and control the start and stop of the synchronization task of the task log of the offline data development.
[0077] The output end of the data integration directory module is connected to the input end of the offline synchronization module; the output end of the offline synchronization module is connected to the input end of the production environment operation and maintenance management module; the output end of the production environment operation and maintenance management module is connected to the input end of the intelligent monitoring module.
[0078] The data integration directory module includes a software development business unit and an offline synchronization task construction unit;
[0079] The software development business unit is used to record the software development business process name to form a software development business list; the offline synchronization task construction unit is used to drag and drop the offline synchronization node to construct an offline synchronization task;
[0080] The output end of the software development business unit is connected to the input end of the offline synchronization task construction unit.
[0081] The offline synchronization module includes an offline synchronization test unit and an offline synchronization configuration unit;
[0082] The offline synchronization test unit is used to configure the synchronization network link under the offline synchronization task, select the data source and data destination of the offline synchronization task, and the resource group for executing the synchronization task, and test the connectivity; the offline synchronization configuration unit is used to edit the script and configure the relevant attributes during the automatic scheduling of the task under the periodically scheduled offline synchronization task. The relevant attributes include: configuring the node scheduling attribute, configuring the time attribute, and configuring the resource attribute;
[0083] The output end of the offline synchronization test unit is connected to the input end of the offline synchronization configuration unit.
[0084] The production environment operation and maintenance management module includes a production environment operation and maintenance center and a large model analysis unit;
[0085] The production environment operation and maintenance center is used to import offline synchronization tasks and call large model processing; the large model analysis unit forms large models under different resource groups based on sample data;
[0086] The output end of the large model analysis unit is connected to the input end of the production environment operation and maintenance center.
[0087] The intelligent monitoring module includes an intelligent monitoring unit and an automatic start-stop control unit;
[0088] The intelligent monitoring unit is used to intelligently monitor the offline data development task log according to the dirty data processing node and the offline synchronization task, and control the start and stop of the synchronization task of the offline data development task log; the automatic start-stop control unit is used to obtain the offline synchronization task and the number of configured dirty data processing nodes under the offline synchronization task, and perform intelligent monitoring of the offline data development task log. If the number of dirty data generated within the period of the configured time attribute is less than or equal to the number of configured dirty data processing nodes, the synchronization task of the offline data development task log is normally carried out; if the number of dirty data generated within the period of the configured time attribute is higher than the number of configured dirty data processing nodes, the synchronization task of the offline data development task log is stopped, the synchronized offline task data is retained, and an error is reported at the administrator port at the same time.
[0089] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.
[0090] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An intelligent monitoring method for the task log of offline data development based on a large model, characterized in that: The method includes the following steps: S1. Obtain the name of the software development business process, drag the offline synchronization node in the data integration directory module, and construct an offline synchronization task; S2. Under the offline synchronization task, configure the synchronization network link, select the data source and data destination of the offline synchronization task, as well as the resource group for executing the synchronization task, test the connectivity, edit the script, and configure the offline synchronization task; S3. Import the offline synchronization task into the production environment operation and maintenance center, and the production environment operation and maintenance center calls the large model to configure the dirty data processing node; S4. According to the dirty data processing node and the offline synchronization task, perform intelligent monitoring of the offline data development task log, and control the start and stop of the synchronization task of the offline data development task log; In step S3, it also includes: Based on different resource groups for executing the synchronization task, large models under different resource groups are respectively constructed. The construction method of the large model under any resource group includes: Construct a sample data set, select the original data under the same resource group, where the original data refers to the manually set data of the user administrator, to form a specific data set [m0, y0]. Here, m0 refers to the offline synchronization task feature group, specifically including the total number of offline synchronization tasks and the number of associations between offline synchronization tasks. The association means that there is data interaction between any two offline synchronization tasks, which is defined as the number of associations once; y0 refers to the number of configured dirty data processing nodes; assign values to m0 in a fixed weight format, including: m0 = z1 * a0 + (1 - z1) * b0 Where, z1 refers to the weight set by the system; a0 and b0 respectively represent the normalized data of the total number of offline synchronization tasks and the number of associations between offline synchronization tasks; Based on the sample data set, construct a frequency histogram. In the frequency histogram, the area of the rectangle is defined as the interval frequency, and the height of the rectangle is the average frequency density of the interval. Then: where x i refers to the sample data i in the sample data set; f refers to the probability density function based on the sample data set; F refers to the cumulative distribution function based on the sample data set; h represents the bandwidth of kernel density estimation, set by the system; F(x i +h), F(x i -h) represent the cumulative distribution function values at (x i +h), (x i -h) respectively. That is, for a continuous function, the sum of the occurrence probabilities of all values less than or equal to (x i +h) is the cumulative distribution function value at (x i +h); the sum of the occurrence probabilities of all values less than or equal to (x i -h) is the cumulative distribution function value at (x i -h); In a two-dimensional scenario of a specific data set [m0, y0], for any new data x, find its closest range [x i -h, x i +h] in the sample data set; and calculate the probability density function for any new data x: Among them, s represents the serial number; n represents the total number of samples in the sample data set; K0 represents the uniform kernel function; dis(x, [[m s , y s ) represents the distance between x and [[m s , y s ; According to the total number of offline synchronization tasks and the number of associations between offline synchronization tasks monitored by the system, form a specific new data set, and thus output the number of configured dirty data processing nodes under the specific new data set.
2. The intelligent monitoring method for the offline data development task log based on the large model according to claim 1, characterized in that: In step S2, it also includes: Select the switch network segment bound by the resource group, the original EIP of the old version resource group itself, or the EIP configured by the new version general-purpose resource group bound to the VPC, and add it to the whitelist of the data source; Construct the data association between the data source and DataWorks, and add the DataWorks data source to test the connectivity. Specifically, it includes: Under the data integration directory module, log in to the DataWorks console, select data integration in the navigation bar, enter data integration after selecting the corresponding workspace, select the data source in the navigation bar, then add a new data source in the data source list, set the new data source as the current data source, and configure the relevant connection parameters to achieve testing of the connectivity; The configuration of the offline synchronization task also includes: For the offline synchronization task with periodic scheduling, relevant attributes for automatic task scheduling need to be configured. The relevant attributes include: configuring node scheduling attributes, configuring time attributes, and configuring resource attributes; The configuration of the node scheduling attributes includes: assigning scheduling parameters to custom variables, and at the same time supporting the assignment of constants; The configured time attribute includes: defining the periodic scheduling method of tasks in the production environment; The configured resource attribute includes: the scheduling resource group used to define the execution resources of data integration tasks in the scenario of offline synchronization tasks with periodic scheduling.
3. The intelligent monitoring method for the task log of offline data development based on a large model according to claim 2, wherein: It also includes: Obtain the offline synchronization task and the configured number of dirty data processing nodes under the offline synchronization task, and conduct intelligent monitoring of the offline data development task log. If the number of dirty data generated within the period under the configured time attribute is less than or equal to the configured number of dirty data processing nodes, the synchronization task of the offline data development task log is normally carried out; if the number of dirty data generated within the period under the configured time attribute is higher than the configured number of dirty data processing nodes, the synchronization task of the offline data development task log is stopped, the synchronized offline task data is retained, and an error is reported at the administrator port at the same time.
4. An intelligent monitoring system for offline data development task logs based on a large model, which is used to implement an intelligent monitoring method for offline data development task logs based on a large model as described in claim 1, and is characterized in that: The system includes: a data integration catalog module, an offline synchronization module, a production environment operation and maintenance management module, and an intelligent monitoring module; The data integration catalog module is used to obtain the software development business process name, drag the offline synchronization node in the data integration catalog module to construct an offline synchronization task; the offline synchronization module is used to configure the synchronization network link under the offline synchronization task, select the data source and data destination of the offline synchronization task, and the resource group for executing the synchronization task, test the connectivity, edit the script, and configure the offline synchronization task; the production environment operation and maintenance management module is used to import the offline synchronization task into the production environment operation and maintenance center, and the production environment operation and maintenance center calls the large model to configure the dirty data processing node; the intelligent monitoring module is used to conduct intelligent monitoring of the offline data development task log according to the dirty data processing node and the offline synchronization task, and control the start and stop of the synchronization task of the offline data development task log. The output end of the data integration catalog module is connected to the input end of the offline synchronization module; the output end of the offline synchronization module is connected to the input end of the production environment operation and maintenance management module; the output end of the production environment operation and maintenance management module is connected to the input end of the intelligent monitoring module.
5. An intelligent monitoring system for the task log of offline data development based on a large model according to claim 4, characterized in that: The data integration catalog module includes a software development business unit and an offline synchronization task construction unit; The software development business unit is used to record the software development business process name to form a software development business list; the offline synchronization task construction unit is used to drag the offline synchronization node to construct an offline synchronization task; The output end of the software development business unit is connected to the input end of the offline synchronization task construction unit.
6. The intelligent monitoring system for the task log of offline data development based on a large model according to claim 5, wherein: The offline synchronization module includes an offline synchronization test unit and an offline synchronization configuration unit; The offline synchronization test unit is used to configure the synchronization network link under the offline synchronization task, select the data source and data destination of the offline synchronization task, and the resource group for executing the synchronization task, and test the connectivity; the offline synchronization configuration unit is used to edit the script and configure the relevant attributes during automatic task scheduling under the offline synchronization task with periodic scheduling. The relevant attributes include: configured node scheduling attribute, configured time attribute, and configured resource attribute; The output end of the offline synchronization test unit is connected to the input end of the offline synchronization configuration unit.
7. An intelligent monitoring system for the task log of offline data development based on a large model according to claim 6, characterized in that: The production environment operation and maintenance management module includes a production environment operation and maintenance center and a large model analysis unit; The production environment operation and maintenance center is used to import offline synchronization tasks and call large model processing; the large model analysis unit forms large models under different resource groups based on sample data; The output end of the large model analysis unit is connected to the input end of the production environment operation and maintenance center.
8. An intelligent monitoring system for the task log of offline data development based on a large model according to claim 7, characterized in that: The intelligent monitoring module includes an intelligent monitoring unit and an automatic start-stop control unit; The intelligent monitoring unit is used to perform intelligent monitoring of the offline data development task log according to the dirty data processing node and the offline synchronization task, and control the start and stop of the synchronization task of the offline data development task log; the automatic start-stop control unit is used to obtain the offline synchronization task and the number of configured dirty data processing nodes under the offline synchronization task, perform intelligent monitoring of the offline data development task log, and if the number of dirty data generated within the period under the configured time attribute is less than or equal to the number of configured dirty data processing nodes, the synchronization task of the offline data development task log is normally carried out; if the number of dirty data generated within the period under the configured time attribute is higher than the number of configured dirty data processing nodes, the synchronization task of the offline data development task log is stopped, the already synchronized offline task data is retained, and an error is reported at the administrator port at the same time.
Citation Information
Patent Citations
Fine adjustment method and device for large model in banking business, equipment and storage medium
CN116579402A
Data management platform taking metadata as core and implementation method
CN116910078A