Data processing method, computer equipment and computer storage medium

Through dynamic programming of data processing topology dependencies and real-time monitoring technology, the problems of poor adaptability and delayed anomaly detection in traditional data processing methods are solved, and efficient, reliable and flexible anomaly monitoring of the data processing process is achieved.

CN120597115APending Publication Date: 2025-09-05祁卓玛
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510760185.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Traditional data processing methods are difficult to adapt to complex business rules and dynamically changing data environments, resulting in high development and maintenance costs. They also lack effective real-time anomaly identification and feedback mechanisms, affecting the stability and reliability of data processing.

Method used

By dynamically planning data processing topology dependencies, quantifying the status of intermediate data processing nodes, and building a data processing process encoding label queue, combined with a deep learning neural network model for real-time monitoring and anomaly identification, intelligent optimization and anomaly detection of data processing processes can be achieved.

Benefits of technology

It improves the flexibility and adaptability of data processing, ensures the accuracy and reliability of the data processing process, reduces downtime and maintenance costs, and enables timely detection and accurate tracing of abnormal situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597115A_ABST
    Figure CN120597115A_ABST
Patent Text Reader

Abstract

The invention relates to a data processing method, computer equipment and a computer storage medium. The method comprises the following steps: acquiring to-be-processed initial data, and obtaining to-be-processed structured data and a data coding label; matching a starting point topological node and an end point topological node of the to-be-processed structured data in the data processing topological model based on the data coding label; data processing topology connection relation information between the starting point topology node and the end point topology node is obtained from the data processing topology model, and data processing path decision information is generated; processing the to-be-processed structured data according to the data processing path decision information, and constructing a data processing flow coding label queue; and inputting the data processing flow coding queue data into the data processing flow state abnormity identification model to obtain a data processing flow state identification result. By adopting the method, the data processing efficiency, the system reliability and the service adaptability can be improved through coding label driven topology path decision and full-flow state monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing, and in particular relates to a data processing method, computer equipment and computer storage medium. Background Art

[0002] With the rapid development of digital technology, data volume and business complexity are exploding, posing numerous challenges to traditional data processing methods. Every day, businesses and organizations must process massive amounts of data from various sources. This data often has diverse formats, structures, and semantics. Directly processing raw data is not only inefficient but can also lead to erroneous results. Currently, the automation and intelligentization of data processing processes has become a significant development trend. An increasing number of businesses and organizations are adopting intelligent data processing systems to improve the efficiency and quality of data processing.

[0003] Traditional data processing methods typically rely on linear processes. However, these fixed processes struggle to adapt to complex business rules and dynamically changing data environments, often resulting in high development and maintenance costs. Furthermore, existing technologies often lack effective real-time identification and feedback mechanisms for abnormalities that may arise during data processing, making it difficult to ensure stable and reliable data processing.

[0004] The prior art, with authorization announcement number CN116450367B, provides a data processing method, apparatus, device, and storage medium that can reduce data calculation errors and improve computational efficiency. However, this technical solution struggles with complex business rule calculation path planning, an imperfect data processing process traceability mechanism, and a single and lagging anomaly detection dimension. Summary of the Invention

[0005] Based on this, it is necessary to provide a data processing method, computer equipment and computer storage medium that can dynamically plan data processing topology dependencies, comprehensively quantify the status of intermediate data processing nodes and monitor the entire data processing process with coded labels to address the above technical problems.

[0006] In a first aspect, the present application provides a data processing method, comprising:

[0007] Obtaining the initial data to be processed, and preprocessing the initial data to be processed to obtain the structured data to be processed and the data encoding label of the structured data to be processed;

[0008] Matching the starting topology node and the ending topology node of the structured data to be processed in the data processing topology model based on the data encoding label;

[0009] Obtaining data processing topology connection relationship information between a starting point topology node and an end point topology node from a data processing topology model, identifying intermediate topology nodes, and generating data processing path decision information;

[0010] Process the structured data to be processed according to the data processing path decision information, and build a data processing process encoding label queue, which is used to represent the processing flow of the structured data to be processed from the starting topological node to the end topological node;

[0011] The data processing flow encoding queue data is input into the data processing flow state abnormality recognition model to obtain the data processing flow state recognition result.

[0012] In one embodiment, constructing a data processing flow encoding tag queue includes:

[0013] Initialize the data processing flow encoding label queue using the data encoding label and the starting node encoding label of the starting topology node as the starting queue data;

[0014] Using the data processing intermediary structured data corresponding to the data processing path decision information and the intermediary node encoding label of the intermediary topology node as the queue adding data, updating the data processing flow encoding label queue;

[0015] The processing result structured data corresponding to the terminal topology node and the terminal node coding label of the terminal topology node are used as the end queue data to complete the construction of the data processing flow coding label queue.

[0016] In one embodiment, the intermediate node encoding tag includes a real-time intermediate node state encoding tag and an intermediate node type encoding tag. The data processing intermediate structured data corresponding to the data processing path decision information and the intermediate node encoding tag of the intermediate topology node are used as queue data to update the data processing flow encoding tag queue, including:

[0017] Based on the intermediary node type coding label, a decision tree model is used to analyze the intermediary node state parameters of the intermediary topology nodes;

[0018] The real-time intermediary node state encoding label is generated based on the intermediary node state parameters.

[0019] In one embodiment, the intermediary node status parameters include a connection capacity status parameter, a processing time status parameter, and a verification result status parameter. The real-time intermediary node status coding tag includes a real-time intermediary node comprehensive status score. The real-time intermediary node comprehensive status score is used to characterize the computing pressure and data anomaly level of the intermediary topology node. The expression for the real-time intermediary node comprehensive status score is:

[0020]

[0021] Where, is the real-time comprehensive status score of the intermediary node of the i-th intermediary topology node, and are the fuzzy gating function term coefficients of the connection capacity state parameter, the processing time state parameter and the verification result state parameter, respectively, of the i-th intermediary topology node, CCS ,λ PTS ,λ VRS and λ P are the connection capacity state parameter weight coefficient, processing time state parameter weight coefficient, verification result state parameter weight coefficient, and connection capacity state parameter combined with processing time state parameter balancing weight coefficient of the i-th intermediary topology node, and They are the connection capacity state parameter, processing time state parameter and verification result state parameter of the i-th intermediary topology node respectively.

[0022] In one embodiment, the intermediate node coding label includes an intermediate node location coding label, the data processing flow state identification result includes an abnormal intermediate topology node identification result, the data processing flow state abnormality identification model includes a data processing flow intermediate node state abnormality identification sub-model, and the data processing flow coding queue data is input into the data processing flow state abnormality identification model to obtain the data processing flow state identification result, including:

[0023] Input the real-time comprehensive status scores of the intermediary nodes in the data processing process encoding queue data into the data processing process intermediary node status anomaly identification sub-model to identify the intermediary topology nodes with anomalies;

[0024] Based on the intermediary node location coding label of the intermediary topology node with abnormality, a historical intermediary node state coding label dataset of the intermediary topology node with abnormality is obtained;

[0025] The abnormal intermediary topology node identification results are generated according to the historical intermediary node state coding labels and the real-time intermediary node state coding labels.

[0026] In one embodiment, generating abnormal intermediary topology node identification results based on a historical intermediary node state encoding label dataset and a real-time intermediary node state encoding label includes:

[0027] Input the historical intermediary node state encoding label dataset into the historical statistical analysis model to generate historical intermediary node state encoding feature labels;

[0028] Obtain the state encoding label segmentation identification information corresponding to the intermediate node type encoding label;

[0029] Based on the state coding label segmentation identification information, the historical intermediary node state coding feature label and the real-time intermediary node state coding label are segmented and paired to obtain the paired historical intermediary node state coding feature sub-label and the real-time intermediary node state coding sub-label;

[0030] Calculating the error between each paired historical intermediate node state encoding feature sub-label and the real-time intermediate node state encoding sub-label to obtain intermediate node state encoding sub-label error data between each paired historical intermediate node state encoding feature sub-label and the real-time intermediate node state encoding sub-label;

[0031] The abnormal intermediary topology node identification results are generated based on the analysis of the intermediary node state encoding sub-label error data.

[0032] In one embodiment, the historical statistical analysis model is a deep learning neural network model constructed based on an improved variational autoencoder combined with a generative adversarial network model.

[0033] In one embodiment, the intermediate node type coding label includes a calculation intermediate node type coding label and a determination intermediate node type coding label, and the determination intermediate node type coding label includes a data diversion determination intermediate node type coding label and a data anomaly determination data intermediate node type coding label.

[0034] In a second aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any method of the first aspect of the present application are implemented.

[0035] In a third aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any method of the first aspect of the present application.

[0036] The aforementioned data processing method, computer device, and computer storage medium achieve data standardization and semantic labeling by preprocessing raw data to obtain structured data. This addresses the difficulty in processing diverse raw data formats in traditional data processing. By obtaining data processing path information based on a topological model, they enable intelligent optimization of the data processing process, dynamically adapting to changes in data type, structure, and requirements. Furthermore, by analyzing the encoded label queues within the data processing process using an anomaly recognition model, they enable real-time monitoring of the data processing process status, enabling timely detection of anomalies such as data loss, processing delays, and quality degradation, thereby improving the stability and accuracy of data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0038] Figure 1 A schematic diagram of an application environment of a data processing method provided in one embodiment of the present application;

[0039] Figure 2 A flowchart of a data processing method provided in one embodiment of the present application;

[0040] Figure 3 A schematic diagram of a data processing topology connection structure of a data processing method provided in one embodiment of the present application;

[0041] Figure 4 A flowchart of another data processing method provided in one embodiment of the present application;

[0042] Figure 5 A flowchart of a method for generating abnormal intermediary topology node identification results provided by one embodiment of the present application;

[0043] Figure 6 A schematic diagram of the structure of a data processing system provided in one embodiment of the present application. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0045] The data processing method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the data acquisition terminal 102 and the data response terminal 103 can communicate with the server 101 through the network. The database 104 can store the data that the server 101 needs to process. The database 104 can be integrated on the server 101, or it can be placed on the cloud or other network servers. The server 101 can collect the original data that needs to be processed through the data acquisition terminal 102, and perform data processing to obtain data processing result information. The server 101 can send the data processing result information to the data response terminal 103, and control the data response terminal 103 to implement the data response operation through the data processing result information. Among them, the data acquisition terminal 102 can be but is not limited to various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The portable wearable devices can be smart watches, smart bracelets and head-mounted devices, etc. The server 101 can be implemented with an independent server or a server cluster consisting of multiple servers.

[0046] In an exemplary embodiment, Figure 3 As shown, a data processing method is provided, which is applied to Figure 2 The server 101 in the example is used for explanation, and the steps include the following steps S301 to S305.

[0047] in:

[0048] Step S201 : obtaining the initial data to be processed, and preprocessing the initial data to be processed to obtain the structured data to be processed and the data encoding label of the structured data to be processed.

[0049] Specifically, the server 101 may acquire the to-be-processed initial data based on the data acquisition terminal, and may pre-process the to-be-processed initial data to obtain the to-be-processed structured data and a data encoding label of the to-be-processed structured data.

[0050] Optionally, the preprocessing performed by the server 101 on the initial data to be processed may include, but is not limited to, data cleaning, format conversion, data labeling, and data integration.

[0051] Step S202 : matching the starting point topological node and the ending point topological node of the structured data to be processed in the data processing topological model based on the data encoding label.

[0052] Specifically, the server 101 may set the data acquisition terminal as the starting topology node 302 in the data processing topology model, and may update the attributes of the starting topology node 302 based on the structured data to be processed and the data encoding tag of the structured data to be processed. The server 101 may identify the ending topology node 303 in the data processing topology model that matches the structured data to be processed based on the data encoding tag of the structured data to be processed and the starting topology node 302.

[0053] Optionally, the server may set the possible data response terminal corresponding to the structured data to be processed as the end point topology node 303 .

[0054] Illustratively, the starting point topological node 302 may include, but is not limited to, a plurality of data acquisition terminals, and the ending point topological node 303 may include, but is not limited to, a plurality of data response terminals.

[0055] Step S203 : obtaining data processing topology connection relationship information between the starting point topology node and the end point topology node from the data processing topology model, identifying the intermediate topology nodes, and generating data processing path decision information.

[0056] Specifically, the server 101 can obtain the data processing topology connection relationship information between the starting topology node and the ending topology node from the data processing topology model based on the identified starting topology node 302 and ending topology node 303 in combination with the data encoding label of the structured data to be processed, identify the intermediate topology node 305, and generate data processing path decision information.

[0057] Optionally, the intermediary topology nodes may include computing intermediary topology nodes and diversion intermediary topology nodes.

[0058] Step S204: Process the structured data to be processed according to the data processing path decision information, and build a data processing flow encoding label queue.

[0059] Specifically, the server 101 may process the structured data to be processed according to the generated data processing path decision information. The server 101 may construct and update the data processing flow encoding label queue based on the location identification information, state identification information, and intermediate data identification information of the starting point topological node 302, the end point topological node 303, and the intermediate topological node 305 in the process of processing the structured data to be processed.

[0060] Optionally, the data processing flow encoding tag queue may be used to represent the processing flow of processing the structured data to be processed from the starting topological node to the ending topological node.

[0061] Step S205 , inputting the data processing flow coding queue data into the data processing flow state abnormality recognition model to obtain the data processing flow state recognition result.

[0062] Specifically, the server 101 may input the generated data processing flow encoding queue data into a data processing flow state anomaly recognition model carried by the server 101 to obtain a data processing flow state recognition result.

[0063] Illustratively, the data processing flow state identification result can be used to represent the state information of the intermediary topology node 305 called in the process of processing the structured data to be processed.

[0064] In the above data processing method, by obtaining the initial data to be processed and performing preprocessing, noise data and missing data can be effectively removed, and by realizing structured management of the data to be processed, the efficiency and accuracy of the data processing process allocation can be improved, thereby improving the stability of data processing; by generating data coding labels in the preprocessing process, it can facilitate subsequent data management and tracking, which can help to further improve the efficiency and accuracy of data processing, and can simplify data retrieval and verification operations in the data processing process; by matching the starting topological node and the end topological node of the structured data to be processed in the data processing topology model based on the data coding labels, the position of the data to be processed in the data processing topology model can be determined more efficiently, ensuring the accuracy and pertinence of the data processing process, thereby helping to improve the reliability of the entire data processing process.

[0065] Furthermore, in the above-mentioned data processing method, by obtaining data processing topology connection relationship information from the data processing topology model, the optimal data processing path can be decided in real time according to the data processing topology model, thereby improving the flexibility and adaptability of data processing, and thus ensuring that the data processing process can adapt to different business scenarios and needs, and improving the scalability of the data processing system; by obtaining the data processing process status identification result based on the data processing process coding queue identification, efficient real-time monitoring of the data processing process can be achieved, thereby reducing the failure time and maintenance cost of the data processing system, and improving the maintainability of the data processing system.

[0066] In an optional embodiment of the present application, please refer to Figure 4 , building a data processing flow encoding label queue may include:

[0067] Step S405 , initializing the data processing flow coding label queue using the data coding label and the starting node coding label of the starting topological node as the starting queue data.

[0068] Step S406 , using the data processing intermediary structured data and the intermediary node coding labels of the intermediary topology nodes corresponding to the data processing path decision information as queue addition data, and updating the data processing flow coding label queue.

[0069] Step S407 : Using the processing result structured data corresponding to the end topology node and the end node coding label of the end topology node as the end queue data, the data processing flow coding label queue is constructed.

[0070] Schematically, the data processing flow encoding tag queue can be used to represent the data flow traces and state information of the topological nodes in the data processing flow from the structured data to be processed to the processing result structured data.

[0071] In the above data processing method, by constructing a detailed data processing process encoding label queue, it is possible to achieve refined monitoring of the data processing process, reflect the real-time data processing steps and the status of the data processing nodes, facilitate real-time monitoring of the data processing system, and thus enhance the flexibility, adaptability and traceability of the data processing process.

[0072] In an optional embodiment of the present application, the intermediate node coding tag may include a real-time intermediate node state coding tag and an intermediate node type coding tag. The data processing intermediate structured data corresponding to the data processing path decision information and the intermediate node coding tag of the intermediate topology node are used as queue data to update the data processing flow coding tag queue, which may include:

[0073] Based on the intermediary node type coding labels, a decision tree model is used to analyze the intermediary node state parameters of the intermediary topology nodes.

[0074] The real-time intermediary node state encoding label is generated based on the intermediary node state parameters.

[0075] Exemplarily, the intermediate node type coding label may include, but is not limited to, a data computing intermediate node type coding label, a data offloading intermediate node type coding label, and a data anomaly intermediate node type coding label.

[0076] Furthermore, data diversion intermediary node type coding labels may include but are not limited to preset condition determination intermediary node type coding labels, data aggregation determination intermediary node type coding labels and data priority determination intermediary node type coding labels; data anomaly intermediary node type coding labels may include but are not limited to data format verification intermediary node type coding labels, data access control determination intermediary node type coding labels and data timeliness determination intermediary node type coding labels.

[0077] In an optional embodiment of the present application, the intermediary node status parameters may include a connection capacity status parameter, a processing time status parameter, and a verification result status parameter. The real-time intermediary node status coding tag may include a real-time intermediary node comprehensive status score. The real-time intermediary node comprehensive status score may be used to characterize the degree of computational pressure and data anomaly of the intermediary topology node. The expression for the real-time intermediary node comprehensive status score may be:

[0078]

[0079] Where, is the real-time comprehensive status score of the intermediary node of the i-th intermediary topology node, and are the fuzzy gating function term coefficients of the connection capacity state parameter, the processing time state parameter and the verification result state parameter, respectively, of the i-th intermediary topology node, CCS ,λ PTS ,λ VRS and λ P are the connection capacity state parameter weight coefficient, processing time state parameter weight coefficient, verification result state parameter weight coefficient, and connection capacity state parameter combined with processing time state parameter balancing weight coefficient of the i-th intermediary topology node, and They are the connection capacity state parameter, processing time state parameter and verification result state parameter of the i-th intermediary topology node respectively.

[0080] Illustratively, the connection capacity state parameter may be used to characterize the occupancy state of the cache of the physical data processing device corresponding to the intermediate topology node for connecting to other physical data processing devices.

[0081] In an optional embodiment of the present application, the intermediate node coding label may include an intermediate node positioning coding label, the data processing flow state identification result includes an abnormal intermediate topology node identification result, and the data processing flow state abnormality identification model may include a data processing flow intermediate node state abnormality identification sub-model. Figure 4 , input the data processing process encoding queue data into the data processing process state abnormality recognition model to obtain the data processing process state recognition result, which may include:

[0082] Step S408: input the real-time intermediary node comprehensive status score in the data processing flow encoding queue data into the data processing flow intermediary node status abnormality identification sub-model to identify abnormal intermediary topology nodes.

[0083] Step S409 : acquiring a historical intermediate node state coding label dataset of the abnormal intermediate topology node based on the intermediate node location coding label of the abnormal intermediate topology node.

[0084] Optionally, the server may assign a weight to each historical intermediate node state encoding label data based on the real-time time interval of each historical intermediate node state encoding label data in the historical intermediate node state encoding label data set, and update the historical intermediate node state encoding label data set.

[0085] Step S410 : generating an abnormal intermediary topology node identification result according to the historical intermediary node state coding label and the real-time intermediary node state coding label.

[0086] In the above data processing method, by introducing the intermediary node positioning coding label into the intermediary node coding label, the intermediary topology nodes that may have problems in the data processing process can be quickly located; by combining the historical intermediary node status coding label data set, a comprehensive and in-depth analysis of the intermediary topology nodes with abnormalities can be achieved, and then historical data can be effectively used to conduct a comprehensive evaluation of abnormal intermediary nodes, achieve accurate tracing of abnormal states, and accurately determine the cause of the abnormality and its impact range.

[0087] In an optional embodiment of the present application, Figure 5 As shown in the figure, the abnormal intermediary topology node identification results are generated based on the historical intermediary node state encoding label dataset and the real-time intermediary node state encoding label, including:

[0088] Step S501: Inputting a historical intermediate node state coding label dataset into a historical statistical analysis model to generate historical intermediate node state coding feature labels.

[0089] Step S502: Acquire state coding label segmentation identification information corresponding to the intermediate node type coding label.

[0090] Optionally, the state coding label segmentation identification information may include, but is not limited to, state coding label field segmentation identification information;

[0091] Step S503 : segment and pair the historical intermediate node state encoding feature label and the real-time intermediate node state encoding label based on the state encoding label segmentation identification information to obtain paired historical intermediate node state encoding feature sub-label and real-time intermediate node state encoding sub-label.

[0092] Schematically, the server can split and disassemble the historical intermediate node state coding feature label and the real-time intermediate node state coding label field by field based on the state coding label field split identification information in the state coding label split identification information and pair them to obtain the paired historical intermediate node state coding feature sub-label and real-time intermediate node state coding sub-label.

[0093] Step S504 , calculating the error between each paired historical intermediate node state encoding feature subtag and the real-time intermediate node state encoding subtag, and obtaining intermediate node state encoding subtag error data between each paired historical intermediate node state encoding feature subtag and the real-time intermediate node state encoding subtag.

[0094] Step S505 : generating abnormal intermediary topology node identification results based on the intermediary node state coding sub-label error data analysis.

[0095] In the above data processing method, by generating historical intermediary node state coding feature labels, it is possible to deeply explore the key features and patterns in historical data and improve the accuracy and timeliness of anomaly identification; by segmenting and pairing historical intermediary node state coding feature labels and real-time intermediary node state coding labels based on segmentation identification information, the comparison between historical data and real-time data can be made more accurate, which is conducive to targeted abnormal state analysis.

[0096] In an optional embodiment of the present application, the historical statistical analysis model may be a deep learning neural network model constructed based on an improved variational autoencoder combined with a generative adversarial network model.

[0097] In an optional embodiment of the present application, the intermediate node type coding label may include a computing intermediate node type coding label and a determining intermediate node type coding label, and the determining intermediate node type coding label includes a data diversion determining intermediate node type coding label and a data anomaly determining data intermediate node type coding label.

[0098] In an exemplary embodiment of the present application, Figure 4 As shown, another data processing method is provided, including the following steps S401 to S410.

[0099] Step S401 : obtaining the initial data to be processed, and preprocessing the initial data to be processed to obtain the structured data to be processed and the data encoding label of the structured data to be processed.

[0100] Step S402 : matching the starting topology node and the ending topology node of the structured data to be processed in the data processing topology model based on the data encoding label.

[0101] Step S403 : obtaining data processing topology connection relationship information between the starting point topology node and the end point topology node from the data processing topology model, identifying the intermediate topology nodes, and generating data processing path decision information.

[0102] Step S404: Process the structured data to be processed according to the data processing path decision information.

[0103] Step S405 , initializing the data processing flow coding label queue using the data coding label and the starting node coding label of the starting topological node as the starting queue data.

[0104] Step S406 , using the data processing intermediary structured data and the intermediary node coding labels of the intermediary topology nodes corresponding to the data processing path decision information as queue addition data, and updating the data processing flow coding label queue.

[0105] Step S407 : Using the processing result structured data corresponding to the end topology node and the end node coding label of the end topology node as the end queue data, the data processing flow coding label queue is constructed.

[0106] Step S408: input the real-time intermediary node comprehensive status score in the data processing flow encoding queue data into the data processing flow intermediary node status abnormality identification sub-model to identify abnormal intermediary topology nodes.

[0107] Step S409 : acquiring a historical intermediate node state coding label dataset of the abnormal intermediate topology node based on the intermediate node location coding label of the abnormal intermediate topology node.

[0108] Step S410 : generating an abnormal intermediary topology node identification result according to the historical intermediary node state coding label and the real-time intermediary node state coding label.

[0109] In the above-mentioned data processing method, through efficient data encoding technology, accurate topological node matching technology, dynamic data processing path decision technology, and real-time status monitoring and anomaly identification technology, it is possible to achieve efficient management, accurate planning and real-time anomaly monitoring of the data processing process, thereby solving technical problems existing in traditional data processing methods such as low quality of data processing results, rigid data processing process, untimely discovery of abnormal status of data processing process, and difficulty in locating and identifying abnormal status of data processing process.

[0110] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0111] Based on the same inventive concept, the present application also provides a data processing system for implementing the aforementioned data processing method. The implementation solution provided by this system is similar to the implementation solution described in the aforementioned method. Therefore, the specific limitations in one or more data processing system embodiments provided below can be found in the above-mentioned limitations on the data processing method and will not be repeated here.

[0112] In an exemplary embodiment, Figure 6 As shown, a data processing system 600 is provided, including:

[0113] The initial data preprocessing module 601 may be used to obtain the initial data to be processed, and preprocess the initial data to be processed to obtain the structured data to be processed and the data encoding label of the structured data to be processed.

[0114] The processing path topology analysis module 602 may be used to match the starting topology node and the ending topology node of the structured data to be processed in the data processing topology model based on the data encoding label.

[0115] The data processing path decision module 603 may be used to obtain data processing topology connection relationship information between a starting topology node and an ending topology node from a data processing topology model, identify intermediate topology nodes, and generate data processing path decision information.

[0116] The process coding label management module 604 can be used to process the structured data to be processed according to the data processing path decision information, and construct a data processing process coding label queue, which is used to represent the processing process of the structured data to be processed from the starting topological node to the end topological node.

[0117] The process abnormal state identification module 605 can be used to input the data processing process encoding queue data into the data processing process state abnormality identification model to obtain the data processing process state identification result.

[0118] In an optional embodiment of the present application, the process code label management module 604 may also be used to:

[0119] The data encoding label and the starting node encoding label of the starting topology node are used as the starting queue data to initialize the data processing flow encoding label queue.

[0120] The data processing intermediary structured data corresponding to the data processing path decision information and the intermediary node encoding label of the intermediary topology node are used as queue adding data to update the data processing flow encoding label queue.

[0121] The processing result structured data corresponding to the terminal topology node and the terminal node coding label of the terminal topology node are used as the end queue data to complete the construction of the data processing flow coding label queue.

[0122] In an optional embodiment of the present application, the process code label management module 604 may also be used to:

[0123] Based on the intermediary node type coding labels, a decision tree model is used to analyze the intermediary node state parameters of the intermediary topology nodes.

[0124] The real-time intermediary node state encoding label is generated based on the intermediary node state parameters.

[0125] In an optional embodiment of the present application, the process abnormal state identification module 605 may also be used to:

[0126] The real-time comprehensive status scores of the intermediary nodes in the data processing process encoding queue data are input into the data processing process intermediary node status anomaly identification sub-model to identify the intermediary topology nodes with anomalies.

[0127] A historical intermediary node state encoding label dataset of the intermediary topology node with abnormalities is obtained based on the intermediary node location encoding label of the intermediary topology node with abnormalities.

[0128] The abnormal intermediary topology node identification results are generated according to the historical intermediary node state coding labels and the real-time intermediary node state coding labels.

[0129] In an optional embodiment of the present application, the process abnormal state identification module 605 may also be used to:

[0130] The historical intermediary node state encoding label dataset is input into the historical statistical analysis model to generate historical intermediary node state encoding feature labels.

[0131] Get the state encoding label segmentation identification information corresponding to the intermediate node type encoding label.

[0132] The historical intermediate node state encoding feature label and the real-time intermediate node state encoding label are segmented and paired based on the state encoding label segmentation identification information to obtain paired historical intermediate node state encoding feature sub-label and real-time intermediate node state encoding sub-label.

[0133] The errors between each paired historical intermediate node state encoding feature sub-label and the real-time intermediate node state encoding sub-label are calculated to obtain intermediate node state encoding sub-label error data of each paired historical intermediate node state encoding feature sub-label and the real-time intermediate node state encoding sub-label.

[0134] The abnormal intermediary topology node identification results are generated based on the analysis of the intermediary node state encoding sub-label error data.

[0135] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the data processing method as described above when executing the computer program.

[0136] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0137] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separated, and the parts displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the disclosed solution. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0138] The above-described embodiments merely represent several implementation methods of the embodiments of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the concept of the embodiments of the present application, and these modifications and improvements fall within the scope of protection of the embodiments of the present application.

Claims

1. A data processing method, characterized in that: The method comprises: Acquire initial data to be processed, and preprocess the initial data to be processed to obtain structured data to be processed and a data encoding label of the structured data to be processed; Matching a starting point topological node and an ending point topological node of the structured data to be processed in a data processing topological model based on the data encoding label; Acquire data processing topology connection relationship information between the starting point topology node and the end point topology node from the data processing topology model, identify intermediate topology nodes, and generate data processing path decision information; Processing the structured data to be processed according to the data processing path decision information, and constructing a data processing process encoding label queue, wherein the data processing process encoding label queue is used to represent the processing flow of the structured data to be processed from the starting topological node to the ending topological node; The data processing flow encoding queue data is input into a data processing flow state abnormality recognition model to obtain a data processing flow state recognition result.

2. The method according to claim 1, characterized in that The construction of the data processing flow encoding label queue includes: Initializing the data processing flow coding label queue using the data coding label and the starting node coding label of the starting topological node as starting queue data; Using the data processing intermediary structured data and the intermediary node coding labels of the intermediary topology nodes corresponding to the data processing path decision information as queue addition data, updating the data processing flow coding label queue; The data processing flow coding label queue is constructed by using the processing result structured data corresponding to the terminal topology node and the terminal node coding label of the terminal topology node as the end queue data.

3. The method according to claim 2, characterized in that The intermediate node coding tag includes a real-time intermediate node state coding tag and an intermediate node type coding tag. The data processing intermediate structured data corresponding to the data processing path decision information and the intermediate node coding tag of the intermediate topology node are used as queue addition data, and the data processing flow coding tag queue is updated, including: Analyzing the intermediary node state parameters of the intermediary topology nodes using a decision tree model based on the intermediary node type coding label; The real-time intermediate node state encoding label is generated by calculation based on the intermediate node state parameter.

4. The method according to claim 3, characterized in that The intermediary node status parameters include a connection capacity status parameter, a processing time status parameter, and a verification result status parameter. The real-time intermediary node status coding tag includes a real-time intermediary node comprehensive status score. The real-time intermediary node comprehensive status score is used to characterize the computing pressure and data anomaly level of the intermediary topology node. The expression for the real-time intermediary node comprehensive status score is: Where, Score the real-time intermediary node comprehensive status of the i-th intermediary topology node, and are the connection capacity state parameter fuzzy gating function item coefficient, processing time state parameter fuzzy gating function item coefficient and verification result state parameter fuzzy gating function item coefficient of the i-th intermediary topology node, respectively, CCS ,λ PTS ,λ VRS and λ P are the connection capacity state parameter weight coefficient, processing time state parameter weight coefficient, verification result state parameter weight coefficient, and connection capacity state parameter combined with processing time state parameter balancing weight coefficient of the i-th intermediate topology node, and They are respectively the connection capacity state parameter, the processing time state parameter and the verification result state parameter of the i-th intermediate topology node.

5. The method according to claim 4, characterized in that The intermediate node coding label includes an intermediate node positioning coding label, the data processing flow state identification result includes an abnormal intermediate topology node identification result, the data processing flow state abnormality identification model includes a data processing flow intermediate node state abnormality identification sub-model, and the data processing flow coding queue data is input into the data processing flow state abnormality identification model to obtain the data processing flow state identification result, including: Inputting the real-time intermediary node comprehensive status score in the data processing flow encoding queue data into the data processing flow intermediary node status anomaly identification sub-model to identify the intermediary topology node with anomalies; Acquiring a historical intermediate node state coding label dataset of the intermediate topology node with the abnormality based on the intermediate node positioning coding label of the intermediate topology node with the abnormality; The abnormal intermediary topology node identification result is generated according to the historical intermediary node state coding label and the real-time intermediary node state coding label.

6. The method according to claim 5, characterized in that The generating of the abnormal intermediary topology node identification result according to the historical intermediary node state coding label dataset and the real-time intermediary node state coding label includes: Inputting the historical intermediate node state coding label data set into the historical statistical analysis model to generate historical intermediate node state coding feature labels; Obtaining state coding label segmentation identification information corresponding to the intermediate node type coding label; Splitting and pairing the historical intermediate node state coding feature label and the real-time intermediate node state coding label based on the state coding label segmentation identification information to obtain a paired historical intermediate node state coding feature sub-label and a real-time intermediate node state coding sub-label; Calculating the error between each pair of the historical intermediate node state encoding feature subtag and the real-time intermediate node state encoding subtag to obtain intermediate node state encoding subtag error data between each pair of the historical intermediate node state encoding feature subtag and the real-time intermediate node state encoding subtag; The abnormal intermediary topology node identification result is generated according to the intermediary node state coding sub-label error data analysis.

7. The method according to claim 6, characterized in that The historical statistical analysis model is a deep learning neural network model constructed based on an improved variational autoencoder combined with a generative adversarial network model.

8. The method according to any one of claims 1 to 7, characterized in that The intermediate node type coding labels include calculation intermediate node type coding labels and determination intermediate node type coding labels, and the determination intermediate node type coding labels include data diversion determination intermediate node type coding labels and data anomaly determination data intermediate node type coding labels.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Data processing methods, apparatus, equipment and storage media

    CN116450367B