Computer data processing system based on machine learning
The machine learning-based data processing system effectively categorizes and prioritizes multi-source heterogeneous data using a three-dimensional model, enhancing processing efficiency and adaptability by optimizing resource allocation and handling delays.
Patent Information
- Application Number
- CN202510792659.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-13
AI Technical Summary
The prior art is difficult to effectively process multi-source heterogeneous data, resulting in low analysis efficiency and inaccurate priority processing.
Using a computer data processing system based on machine learning, the data multi-dimensional classification module is used to divide the security levels and build a three-dimensional classification model, combining type feature analysis, priority scoring and abnormal monitoring modules to achieve efficient processing and dynamic optimization of multi-source heterogeneous data.
It realizes efficient processing and dynamic optimization of multi-source heterogeneous data, ensures reasonable classification and priority allocation of data, improves the accuracy and efficiency of data processing, and enhances the adaptability and stability of the system.
Smart Images

Figure CN120316587A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a computer data processing system based on machine learning. Background Art
[0002] With the development of information technology, a large amount of multi-source heterogeneous data has been generated in fields such as manufacturing. These data have different structures and sources, so it is necessary to scientifically allocate resources, optimize the processing process, and determine reasonable processing priorities.
[0003] Chinese Patent Publication No.: CN104798043A discloses a data processing method and a computer system, including discretizing a data sample to obtain a data sample in matrix form (S101), training the data sample in matrix form according to a preset classification method to obtain a classification rule set (S102), and converting the classification rule set into a classification rule set recognizable by a data decision platform (S103) and then providing it to the data decision platform (S104), so that the data decision platform can make data decisions according to the classification rule set recognizable by the computer system after conversion. The above process from training the data sample to applying the obtained classification rule set to data decision-making is automatically completed by the computer system, avoiding manual participation. When the data sample changes or the original classification rule set needs to be updated, an updated classification rule set can be obtained in a timely manner. This invention realizes the classification processing training of sample data, but does not realize the comprehensive priority processing analysis of multi-source heterogeneous data, and there are problems of low analysis efficiency for multi-source heterogeneous data and inaccurate priority processing for multi-source heterogeneous data. Summary of the Invention
[0004] The purpose of the present invention is to provide a computer data processing system based on machine learning to solve at least one of the problems existing in the prior art.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A computer data processing system based on machine learning, including: A data multi-dimensional classification module for dividing the security levels of the data content of multi-source heterogeneous data and constructing a three-dimensional classification model based on the real-time nature, data volume, and security level of the multi-source heterogeneous data to determine the comprehensive data type; A type feature analysis module for performing feature analysis on the comprehensive data type, timestamp number, and multi-source heterogeneous data to obtain data features; A priority scoring module for performing weighted scoring on the real-time nature, data volume, security level, and data features of the multi-source heterogeneous data to obtain a priority score, and also for performing matching analysis on the priority score and the comprehensive data type to update the three-dimensional classification model.
[0006] Furthermore, the system further includes: A multi-source data acquisition module for real-time acquisition of multi-source heterogeneous data and source identifiers; A data timestamp construction module for numbering multi-source heterogeneous data according to the acquisition time to obtain a timestamp number; A multi-source data processing module for processing multi-source heterogeneous data according to the data comprehensive type and priority scoring order; A processing exception monitoring module for counting the number of exception processes of multi-source heterogeneous data and performing comparison analysis on the number of exception processes to trigger rule optimization of the three-dimensional classification model.
[0007] Furthermore, the data multi-dimensional classification module is provided with a security level division unit for dividing the security level according to the data content of the multi-source heterogeneous data. The security level division unit divides the security level of the multi-source heterogeneous data into three levels according to different data contents; The data multi-dimensional classification module is also provided with a data volume classification unit for performing comparison analysis on the data volume of the multi-source heterogeneous data to divide the data volume level. The data volume classification unit divides the data volume level of the multi-source heterogeneous data into three levels according to the size of the data volume; The data multi-dimensional classification module is also provided with a time sensitivity analysis unit for dividing the delay tolerance level according to the source identifier. The time sensitivity analysis unit divides the delay tolerance level of the multi-source heterogeneous data into three levels according to different source identifiers.
[0008] Furthermore, the data multi-dimensional classification module is also provided with a three-dimensional classification construction unit for encoding the delay tolerance level, data volume level, and security level to obtain a digital level, and quantitatively representing the delay tolerance level, data volume level, and security level; The three-dimensional classification construction unit analyzes type judgment parameters according to the digital levels of the delay tolerance level, data volume level, and security level of the multi-source heterogeneous data to obtain a type judgment parameter S, and performs comparison analysis on the type judgment parameter to judge the data comprehensive type. The data comprehensive type of the multi-source heterogeneous data is divided into four categories according to the size of the type judgment parameter. The data comprehensive types include emergency security affairs, emergency standard affairs, regular standard affairs, and batch processing affairs.
[0009] Furthermore, the type feature analysis module performs feature analysis on the data comprehensive type, timestamp number, and multi-source heterogeneous data to obtain a data feature L(d).
[0010] Furthermore, the priority scoring module is provided with a data scoring analysis unit, which is used to analyze the level scores according to the delay tolerance level, data volume level and security level of multi-source heterogeneous data, so as to obtain the security level score B1, the data volume level score B2 and the delay tolerance level B3; The priority scoring module is also provided with a scoring weighting analysis unit, which is used to perform weighted analysis on the level scores and data characteristics to obtain the priority score. Let the priority score be Z. If L(d) > θ, set Z = β1×B1 + β2×B2 + β3×B3; otherwise, set Z = [β1 + L(d) / 10]×B1 + [β2 - L(d) / 5]×B2 + [β3 + L(d) / 10]×B3, where β1 represents the first weight, β2 represents the second weight, β3 represents the third weight, β1 + β2 + β3 = 1 and β1×β2×β3 ≠ 0.
[0011] Furthermore, the priority scoring module is also provided with a classification matching and updating unit, which is used to perform matching analysis on the priority score and the data comprehensive type, extract the priority scores corresponding to different data comprehensive types as the comprehensive type scores, and respectively arrange the comprehensive type scores corresponding to each data comprehensive type in ascending order. Analyze the type discrete parameter according to the arranged comprehensive type scores to obtain the type discrete parameter R(m); The classification matching and updating unit performs comparison and analysis on the type discrete parameter to update the three-dimensional classification model. When R(m) > r, add a priority mark to the multi-source heterogeneous data that satisfies Z ≤ Z 0.5×NU(m) + Min(Z∈U(m)), where r represents the update matching threshold.
[0012] Furthermore, the multi-source data processing module responds to and processes multi-source heterogeneous data according to the data comprehensive type and the priority score order. The multi-source data processing module respectively sorts the multi-source heterogeneous data with the same data comprehensive type and with and without priority marks in descending order of the priority score, and preferentially processes the multi-source heterogeneous data with priority marks in the arranged order, and then processes the multi-source heterogeneous data without priority marks in the arranged order. If Z > z1, the multi-source data processing module allocates the multi-source heterogeneous data to the GPU acceleration queue for processing; if Z < z2, the multi-source data processing module classifies the multi-source heterogeneous data into the night batch task pool for processing; if z2 ≤ Z ≤ z1, the multi-source data processing module normally processes the multi-source heterogeneous data; where z1 represents the first processing threshold and z2 represents the second processing threshold.
[0013] Further, the processing anomaly monitoring module is provided with a timeout processing judgment unit, which is used to judge the completion type of multi-source heterogeneous data according to the processing duration, timestamp number, and delay tolerance level. When the delay tolerance level is level one, if T≤γ1, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is normal completion; if T>γ1, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is timeout completion. If the processing duration is a missing value, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is incomplete; when the delay tolerance level is level two, if T≤γ2, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is normal completion; if T>γ2, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is timeout completion. If the processing duration is a missing value, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is incomplete; when the delay tolerance level is level three, the timeout processing judgment unit does not judge the completion type; where T represents the processing duration, γ1 represents the first delay threshold, and γ2 represents the second delay threshold.
[0014] Further, the processing anomaly monitoring module is also provided with an anomaly processing analysis unit, which is used to count the number of times the completion type is timeout completion as the timeout processing times, count the number of times the completion type is incomplete as the incomplete times, and take the sum of 70% of the timeout processing times and the incomplete times as the anomaly processing times; When the number of timeout processing times of the multi-source heterogeneous data of the same data comprehensive type in one second by the anomaly processing analysis unit is greater than or equal to k times, it triggers the rule optimization of the three-dimensional classification model. The rule optimization of the three-dimensional classification model is a weighted analysis process that optimizes the level scores according to the number of timeout processing times of the multi-source heterogeneous data of the same data comprehensive type in one second, optimizes the weights of each level score, and correlates the weights of each level score with the logarithm of the timeout processing times to optimize the three-dimensional classification model.
[0015] The beneficial effects of the present invention are as follows: Through the collaborative work of each unit, the efficient processing and dynamic optimization of multi-source heterogeneous data are realized. The security level division, data volume grading, and time sensitivity analysis ensure the reasonable classification and priority allocation of data; the three-dimensional classification construction and scoring weighted analysis improve the accuracy and efficiency of data processing; the anomaly monitoring and rule optimization mechanism enhance the adaptive ability and stability of the system and improve the data processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 This is a schematic structural diagram of a computer data processing system based on machine learning according to this embodiment.
[0018] Figure 2 This is a schematic structural diagram of the data multi-dimensional classification module according to this embodiment.
[0019] Figure 3 This is a schematic structural diagram of the priority scoring module according to this embodiment.
[0020] Figure 4 This is a schematic structural diagram of the processing anomaly monitoring module according to this embodiment. Detailed implementation manners
[0021] To more clearly illustrate the present invention, the present invention will be further described below in conjunction with preferred embodiments and the accompanying drawings. Similar components in the drawings are denoted by the same reference numerals. Those skilled in the art should understand that the content specifically described below is illustrative rather than restrictive, and should not be used to limit the protection scope of the present invention.
[0022] It should be noted that although terms such as first, second, and third may be used in the embodiments of the present application for description, these descriptions should not be limited to these terms. These terms are only used to distinguish the descriptions. For example, without departing from the scope of the embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first.
[0023] Please refer to Figure 1 as shown. It is a computer data processing system based on machine learning according to this embodiment, including: A multi-source data acquisition module for real-time acquisition of multi-source heterogeneous data and source identifiers. The multi-source heterogeneous data includes data volume and data content. The data content includes, but is not limited to, log files, media data, and document data, etc. The source identifiers include real-time interaction, online Q&A, and offline analysis. The real-time interaction indicates that the multi-source heterogeneous data is data from a high-frequency access interaction source, such as financial transactions and monitoring device sensor data, etc. The online Q&A indicates that the multi-source heterogeneous data is data from an online interaction transmission source, such as hospital online diagnosis and social media chat, etc. The offline analysis indicates that the unit heterogeneous data is data from an offline storage analysis source, such as enterprise database content sorting and preprocessing analysis of stored data, etc. The acquisition method of the multi-source heterogeneous data and source identifiers is user interaction upload.
[0024] Please continue to refer to Figure 1 as shown. The computer data processing system based on machine learning further includes: A data timestamp construction module is used to number multi-source heterogeneous data according to the acquisition time to obtain a timestamp number. The data timestamp construction module is connected to the multi-source data acquisition module.
[0025] Specifically, in this embodiment, the data timestamp construction module numbers the days corresponding to the acquisition time in chronological order to obtain a day number, and numbers each second of a day in chronological order with a second precision to obtain a second number. Then, the day number and the second number are combined as the timestamp number. Let the timestamp number be (d, t), where d represents the day number, d ∈ N + , and t represents the second number, t ∈ N + and t ≤ 86400.
[0026] Specifically, in this embodiment, the data timestamp construction module analyzes the timing number to ensure the timing and continuity of data processing, avoid data conflicts, and improve the standardization of the processing process.
[0027] Please continue to refer to Figure 1 As shown, the computer data processing system based on machine learning further includes: A data multi-dimensional classification module is used to divide the security levels of the data content of multi-source heterogeneous data, and construct a three-dimensional classification model based on the real-time nature, data volume, and security level of the multi-source heterogeneous data to determine the comprehensive data type. The data multi-dimensional classification module is connected to the data timestamp construction module.
[0028] Please refer to Figure 2 As shown, the data multi-dimensional classification module includes: A security level division unit is used to divide the security level according to the data content of multi-source heterogeneous data.
[0029] Specifically, in this embodiment, the security level division unit divides the security level according to the data content of multi-source heterogeneous data. The security level division unit sets the security level of data content of the private type as level one, the security level of data content of the verification access type as level two, and the security level of data content of the public type as level three.
[0030] Specifically, in this embodiment, the data content of the public type data includes, but is not limited to, data publicly available to the public such as media information and tender information, etc. The data content of the verification access type data includes, but is not limited to, data that requires user login verification to access such as server log data and enterprise document data, etc. The data content of the private type data includes, but is not limited to, data related to user personal privacy such as user biometric data, etc.
[0031] It can be understood that in this embodiment, the process of dividing the security level is not specifically limited, and those skilled in the art can freely set it, as long as it meets the classification of the security levels of multi-source heterogeneous data according to different data content types. For example, it can also be set to divide the security level according to the data transmission method.
[0032] Specifically, in this embodiment, the security level division unit divides the security level according to the data content to ensure that sensitive data is protected preferentially, reduce the risk of data leakage, and at the same time optimize resource allocation to avoid overprocessing of data with a low security level.
[0033] Please continue to refer to Figure 2 As shown, the data multi-dimensional classification module further includes: A data volume grading unit for grading the data volume according to the data volume of multi-source heterogeneous data.
[0034] Specifically, in this embodiment, the data volume grading unit performs a comparison and analysis on the data volume of multi-source heterogeneous data to divide the data volume level. If V ≤ v1, the data volume grading unit sets the data volume level of the currently analyzed multi-source heterogeneous data as level one. If v1 < V ≤ v2, the data volume grading unit sets the data volume level of the currently analyzed multi-source heterogeneous data as level two. If V > v2, the data volume grading unit sets the data volume level of the currently analyzed multi-source heterogeneous data as level three. Herein, V represents the data volume of multi-source heterogeneous data, v1 represents the first data volume threshold, 1 ≤ v1 ≤ 2, v2 represents the second data volume threshold, 9 ≤ v2 ≤ 10. It can be understood that in this embodiment, the specific values of the data volume thresholds are not specifically limited, and those skilled in the art can freely set them, as long as the division of the data volume level is met. The optimal values of the data volume thresholds are: v1 = 1, v2 = 10. The unit of the data volume of the multi-source heterogeneous data and the data volume thresholds is GB.
[0035] Specifically, in this embodiment, the data volume grading unit divides the level according to the data volume size to help the system reasonably allocate computing resources, avoid large data volume tasks from occupying too many resources, and improve the overall processing efficiency of the system.
[0036] Please continue to refer to Figure 2 As shown, the data multi-dimensional classification module further includes: A time sensitivity analysis unit for grading the delay tolerance according to the source identifier.
[0037] Specifically, in this embodiment, the time sensitivity analysis unit divides the delay tolerance levels according to the source identifier. The time sensitivity analysis unit sets the delay tolerance level of the multi-source heterogeneous data with the source identifier of real-time interaction to level one; sets the delay tolerance level of the multi-source heterogeneous data with the source identifier of online Q&A to level two; and sets the delay tolerance level of the multi-source heterogeneous data with the source identifier of offline analysis to level three.
[0038] It can be understood that in this embodiment, there is no specific limitation on the division of the delay tolerance levels. Those skilled in the art can freely set them, as long as the division of the maximum delay for processing data from different sources is satisfied, and the tolerance degree of the processing delay time is represented by the delay tolerance levels.
[0039] Specifically, in this embodiment, the time sensitivity analysis unit divides the delay tolerance levels according to the data source identifier to ensure that the data with high real-time requirements is processed first, reduce the delay, and improve the system response speed.
[0040] Please continue to refer to Figure 2 as shown, the data multi-dimensional classification module further includes: A three-dimensional classification construction unit, which is used to construct a three-dimensional classification model according to the delay tolerance level, data volume level, and security level of the multi-source heterogeneous data to judge the comprehensive data type. The three-dimensional classification construction unit is connected to the security level division unit, the data volume grading unit, and the time sensitivity analysis unit.
[0041] Specifically, in this embodiment, the three-dimensional classification construction unit encodes the delay tolerance level, data volume level, and security level to obtain a digital level, and quantitatively represents the delay tolerance level, data volume level, and security level.
[0042] Specifically, in this embodiment, when encoding the delay tolerance level, data volume level, and security level, 1, 2, and 3 are used to represent level one, level two, and level three respectively.
[0043] Specifically, in this embodiment, the three-dimensional classification construction unit analyzes the type judgment parameters according to the digital level analysis types of the delay tolerance level, data volume level, and security level of multi-source heterogeneous data, sets the type judgment parameter as S, sets S = s1 × s2 × s3 × ln[s1 + s2 + s3], and performs a comparison and analysis on the type judgment parameter to determine the comprehensive data type. If S < α1, the three-dimensional classification construction unit determines that the comprehensive data type is an emergency security transaction; if α1 ≤ S < α2, the three-dimensional classification construction unit determines that the comprehensive data type is an emergency standard transaction; if α2 ≤ S < α3, the three-dimensional classification construction unit determines that the comprehensive data type is a regular standard transaction; if S ≥ α3, the three-dimensional classification construction unit determines that the comprehensive data type is a batch processing transaction. Among them, α1 represents the first type threshold, 1.5 ≤ α1 ≤ 2, α2 represents the second type threshold, 3 ≤ α2 ≤ 10, α3 represents the third type threshold, 12 ≤ α3 ≤ 15, and α4 represents the fourth type threshold, 20 ≤ α4 ≤ 30. It can be understood that in this embodiment, the specific values of the type thresholds are not limited, and those skilled in the art can freely set them as long as they meet the judgment of the comprehensive data type. The optimal values of the type thresholds are: α1 = 2, α2 = 7, α3 = 15, and α4 = 25.
[0044] Specifically, in this embodiment, the three-dimensional classification construction unit quantifies the delay tolerance level, data volume level, and security level through digital coding to construct a three-dimensional classification model, accurately judge the comprehensive data type, and provide a reliable basis for subsequent priority scoring.
[0045] Please continue to refer to Figure 1 as shown, the computer data processing system based on machine learning further includes: A type feature analysis module for performing feature analysis on the comprehensive data type, timestamp number, and multi-source heterogeneous data to obtain data features. The type feature analysis module is connected to the data multi-dimensional classification module.
[0046] Specifically, in this embodiment, the type feature analysis module performs feature analysis on the comprehensive data type, timestamp number, and multi-source heterogeneous data to obtain data features, sets the data feature as L(d), and sets , where H(x) represents a feature function, H(x) = sin(x / d) × S × V / [Max(S) × Max(V)], x represents a time variable parameter, x = d × t × 2π / 86400, and Max() represents the maximum value of the data in the parentheses.
[0047] Specifically, in this embodiment, the type feature analysis module combines the data features with the timestamp to generate a feature vector, enhances the parsing ability of the machine learning model for data attributes, and optimizes the classification accuracy.
[0048] Please continue to refer to Figure 1 As shown, the machine learning-based computer data processing system further includes: A priority scoring module for weighted scoring of the real-time performance, data volume, security level, and data characteristics of multi-source heterogeneous data to obtain a priority score, and also for performing matching analysis on the priority score and the data comprehensive type to update the three-dimensional classification model. The priority scoring module is connected to the type feature analysis module.
[0049] Please refer to Figure 3 As shown, the priority scoring module includes: A data scoring analysis unit for analyzing the level score according to the delay tolerance level, data volume level, and security level of multi-source heterogeneous data.
[0050] Specifically, in this embodiment, the data scoring analysis unit analyzes the level score according to the delay tolerance level, data volume level, and security level of multi-source heterogeneous data. The security level score is set as B1, and it is set that B1 = 100 - b1×(n - 1). The data volume level score is set as B2, and it is set that B2 = 100 - Min(Max(lg(V + 1), 0), 3) / 3×100. The delay tolerance level is set as B3, and it is set that B3 = 100 - b2×(n - 1), where b1 represents the first score parameter, 40 ≤ b1 ≤ 50, b2 represents the second score parameter, 20 ≤ b2 ≤ 30, n represents the digital level of the delay tolerance level, data volume level, and security level, n ∈ {1, 2, 3}, and Min() represents extracting the minimum value of the data in the parentheses. It can be understood that in this embodiment, the specific values of the score parameters are not specifically limited, and those skilled in the art can freely set them as long as the analysis of the level score is satisfied. The optimal values of the score parameters are: b1 = 50, b2 = 30.
[0051] Specifically, in this embodiment, the data scoring analysis unit generates a level score according to the delay tolerance level, data volume level, and security level, providing a quantitative basis for the priority score and ensuring the scientific and reasonable scoring process.
[0052] Please continue to refer to Figure 3 As shown, the priority scoring module further includes: A scoring weighted analysis unit for weighted analysis of the level score and data characteristics to obtain a priority score. The scoring weighted analysis unit is connected to the data scoring analysis unit.
[0053] Specifically, in this embodiment, the scoring weighted analysis unit performs weighted analysis on the level scores and data characteristics to obtain a priority score. Set the priority score as Z. If L(d) > θ, set Z = β1×B1 + β2×B2 + β3×B3. Conversely, set Z = [β1 + L(d) / 10]×B1 + [β2 - L(d) / 5]×B2 + [β3 + L(d) / 10]×B3, where β1 represents the first weight, β2 represents the second weight, β3 represents the third weight, β1 + β2 + β3 = 1 and β1×β2×β3 ≠ 0, θ represents the feature threshold, and 0.5 ≤ θ ≤ 0.7. It can be understood that in this embodiment, the value of the feature threshold is not specifically limited, and those skilled in the art can freely set it as long as it meets the analysis of the priority score. The optimal value of the feature threshold is: θ = 0.6.
[0054] Specifically, in this embodiment, the initial values of the weights are set as β1 = 0.3, β2 = 0.3, and β3 = 0.4.
[0055] Specifically, in this embodiment, the scoring weighted analysis unit combines the data characteristics with the level scores for weighted analysis to dynamically generate a priority score, ensuring that high-priority tasks are processed in a timely manner and improving the overall efficiency of the system.
[0056] Please continue to refer to Figure 3 as shown, the priority scoring module further includes: A classification matching and updating unit for performing matching analysis on the priority score and the data comprehensive type to update the three-dimensional classification model. The classification matching and updating unit is connected to the scoring weighted analysis unit.
[0057] Specifically, in this embodiment, the classification matching and updating unit performs matching analysis on the priority score and the data comprehensive type, extracts the priority scores corresponding to different data comprehensive types as the comprehensive type scores, and arranges the comprehensive type scores corresponding to each data comprehensive type in ascending order. Analyze the type dispersion parameter according to the arranged comprehensive type scores. Set the type dispersion parameter as R(m), and set R(m) = (Z 0.75×NU(m) -Z 0.25×NU(m) ) / Min(Max(Z∈U(m)),Min(Z∈U(m + 1))), where m represents the data comprehensive type parameter, m ∈ {1, 2, 3, 4}, m = 1 represents the data comprehensive type as emergency safety affairs, m = 2 represents the data comprehensive type as emergency standard affairs, m = 3 represents regular standard affairs, m = 4 represents batch processing affairs, U(m) represents the set of comprehensive type scores corresponding to each comprehensive data type, NU(m) represents the number of comprehensive type scores corresponding to each comprehensive data, Z 0.75×NU(m) represents the comprehensive type score at the 75% position in the arranged comprehensive type scores, Z0.25×NU(m) It represents the comprehensive type score at the 25th percentile in the sorted comprehensive type scores.
[0058] Specifically, in this embodiment, when analyzing the type discrete parameter, if m = 4, the comprehensive type score in the set U(5) is equal to 0.
[0059] Specifically, in this embodiment, the classification matching and updating unit performs a comparison and analysis on the type discrete parameter to update the three-dimensional classification model. If R(m) ≤ r, the classification matching and updating unit does not update the three-dimensional classification model. If R(m) > r, the classification matching and updating unit adds a priority mark to the multi-source heterogeneous data that satisfies S' and Z ≤ Z 0.5×NU(m) + Min(Z ∈ U(m)), where S' represents the original comparison and analysis process, and Z 0.5×NU(m) represents the comprehensive type score at the 50th percentile in the sorted comprehensive type scores, r represents the update matching threshold, and 0.5 ≤ r ≤ 0.6. It can be understood that in this embodiment, the value of the update matching threshold is not specifically limited, and those skilled in the art can freely set it as long as the update of the three-dimensional classification model is satisfied. The optimal value of the update matching threshold is: r = 0.55.
[0060] Specifically, in the update of the comparison and analysis process of the data comprehensive type in this embodiment, if the data comprehensive type is an emergency security matter, the updated comparison process is to retain the original comparison process and add a priority mark to the multi-source heterogeneous data where S < α1 and Z ≤ Z 0.5×NU(1) + Min(Z ∈ U(1)); if the data comprehensive type is an emergency standard matter, the updated comparison process is to retain the original comparison process and add a priority mark to the multi-source heterogeneous data where α1 ≤ S < α2 and Z ≤ Z 0.5×NU(2) + Min(Z ∈ U(2)).
[0061] Specifically, in this embodiment, the classification matching and updating unit matches the priority score with the data comprehensive type to dynamically update the three-dimensional classification model, ensuring that the system can adjust the processing strategy according to real-time data changes and improve the adaptive ability of the system.
[0062] Please continue to refer to Figure 1 as shown, the computer data processing system based on machine learning further includes: A multi-source data processing module for responding to and processing multi-source heterogeneous data according to the data comprehensive type and the priority score order. The multi-source data processing module is connected to the priority score module.
[0063] Specifically, in this embodiment, the multi-source data processing module responds to and processes multi-source heterogeneous data according to the data integration type and the priority scoring order. The multi-source data processing module sorts the multi-source heterogeneous data with the same data integration type and with and without priority tags in descending order of priority score, and preferentially processes the multi-source heterogeneous data with priority tags in the arranged order, and then processes the multi-source heterogeneous data without priority tags in the arranged order. If Z > z1, the multi-source data processing module allocates the multi-source heterogeneous data to the GPU acceleration queue for processing. If Z < z2, the multi-source data processing module classifies the multi-source heterogeneous data into the night batch task pool for processing. If z2 ≤ Z ≤ z1, the multi-source data processing module processes the multi-source heterogeneous data normally; where z1 represents the first processing threshold, 90 ≤ z1 ≤ 95, z2 represents the second processing threshold, 40 ≤ z2 ≤ 60. It can be understood that in this embodiment, the specific values of the processing thresholds are not limited, and those skilled in the art can freely set them as long as they meet the processing of multi-source heterogeneous data. The optimal values of the processing thresholds are: z1 = 92, z2 = 50.
[0064] Please continue to refer to Figure 1 as shown, the machine learning-based computer data processing system further includes: A processing anomaly monitoring module, which is used to count the number of abnormal processing times of multi-source heterogeneous data, and perform comparison and analysis on the number of abnormal processing times to trigger the rule optimization of the three-dimensional classification model. The processing anomaly monitoring module is connected to the multi-source data processing module.
[0065] Please refer to Figure 4 as shown, the processing anomaly monitoring module includes: A processing duration acquisition unit, which is used to acquire the processing duration of multi-source heterogeneous data. The processing duration of multi-source heterogeneous data is the time interval between the application time and the processing completion time of the multi-source heterogeneous data, and its acquisition method is obtained through data import of the system operation management platform; A timeout processing judgment unit, which is used to judge the completion type of multi-source heterogeneous data according to the processing duration and the delay tolerance level. The timeout processing judgment unit is connected to the processing duration acquisition unit.
[0066] Specifically, in this embodiment, the timeout processing judgment unit determines the completion type of multi-source heterogeneous data based on the processing duration, timestamp number, and delay tolerance level. When the delay tolerance level is level one, if T ≤ γ1, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is normal completion; if T > γ1, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is timeout completion. If the processing duration is a missing value, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is incomplete. When the delay tolerance level is level two, if T ≤ γ2, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is normal completion; if T > γ2, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is timeout completion. If the processing duration is a missing value, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is incomplete. When the delay tolerance level is level three, the timeout processing judgment unit does not judge the completion type. Among them, T represents the processing duration, γ1 represents the first delay threshold, 0.05 ≤ γ1 ≤ 0.07, γ2 represents the second delay threshold, 1.5 ≤ γ2 ≤ 2.2. It can be understood that in this embodiment, the specific values of the delay thresholds are not limited, and those skilled in the art can freely set them as long as they meet the judgment of the completion type. The optimal values of the delay thresholds are γ1 = 0.05 and γ2 = 2. The units of the delay threshold and the processing duration are seconds.
[0067] Specifically, in this embodiment, the timeout processing judgment unit determines the data completion type based on the processing duration and the delay tolerance level, so as to ensure that the system can identify timeout or incomplete tasks and provide a basis for exception handling.
[0068] Please continue to refer to Figure 4 As shown, the processing exception monitoring module further includes: An exception handling analysis unit, which is used to count the number of exception handling times according to the completion type of the multi-source heterogeneous data, and perform comparison and analysis on the number of exception handling times to trigger the rule optimization of the three-dimensional classification model. The exception handling analysis unit is connected to the timeout processing judgment unit.
[0069] Specifically, in this embodiment, the exception handling analysis unit counts the number of times with the completion type of timeout completion as the number of timeout processing times, counts the number of times with the completion type of incomplete as the number of incomplete times, and takes the sum of 70% of the number of timeout processing times and the number of incomplete times as the number of exception handling times.
[0070] Specifically, in this embodiment, when the number of timeout processing times of multi-source heterogeneous data of the same data integration type in one second by the exception handling analysis unit is greater than or equal to k times, the rule optimization of the three-dimensional classification model is triggered. The rule optimization of the three-dimensional classification model is a weighted analysis process for optimizing the level scores according to the number of timeout processing times of multi-source heterogeneous data of the same data integration type in one second. The weights of each level score are optimized, and the weights of each level score are related to the logarithm of the number of timeout processing times to optimize the three-dimensional classification model. The optimized first weight is β1', and it is set that β1' = β1 - lnK k / 2, the optimized second weight is β2', and it is set that β2' = β2 - lnK k / 2, the optimized third weight is β3', and it is set that β3' = β3 + lnK k ; k represents the processing times threshold, 3 ≤ k ≤ 5, and K represents the number of timeout processing times of multi-source heterogeneous data of the same data integration type in one second. It can be understood that in this embodiment, no specific limitation is made on the value of the processing times threshold, and those skilled in the art can freely set it as long as it meets the trigger for the rule optimization of the three-dimensional classification model. The best value of the processing times threshold is: k = 3.
[0071] Specifically, in this embodiment, the exception handling analysis unit counts the number of exception handling times and triggers rule optimization, dynamically adjusts the weight allocation of the classification model, improves the system's ability to handle abnormal situations, and ensures the long-term stable operation of the system.
[0072] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is impossible to list all the implementation manners here. Any obvious changes or modifications derived from the technical solutions of the present invention still fall within the protection scope of the present invention.
Claims
1. A computer data processing system based on machine learning, characterized in that, Including: A data multi-dimensional classification module, which is used to divide the security levels of the data content of multi-source heterogeneous data, and construct a three-dimensional classification model based on the real-time nature, data volume, and security level of the multi-source heterogeneous data to determine the comprehensive data type; A type feature analysis module, which is used to perform feature analysis on the comprehensive data type, timestamp number, and multi-source heterogeneous data to obtain data features; A priority scoring module, which is used to perform weighted scoring on the real-time nature, data volume, security level, and data features of the multi-source heterogeneous data to obtain a priority score, and is also used to perform matching analysis on the priority score and the comprehensive data type to update the three-dimensional classification model.
2. The computer data processing system based on machine learning according to claim 1, characterized in that The system further includes: A multi-source data acquisition module, which is used to acquire multi-source heterogeneous data and source identifiers in real time; A data timestamp construction module, which is used to number the multi-source heterogeneous data according to the acquisition time to obtain a timestamp number; A multi-source data processing module, which is used to respond to and process multi-source heterogeneous data according to the comprehensive data type and priority score order; A processing anomaly monitoring module, which is used to count the number of anomaly processes of the multi-source heterogeneous data, and perform comparison analysis on the number of anomaly processes to trigger rule optimization of the three-dimensional classification model.
3. The computer data processing system based on machine learning according to claim 1, wherein The data multi-dimensional classification module is provided with a security level division unit, which is used to divide the security level according to the data content of the multi-source heterogeneous data. The security level division unit divides the security level of the multi-source heterogeneous data into three levels according to different data contents; The data multi-dimensional classification module is further provided with a data volume grading unit, which is used to perform comparison analysis on the data volume of the multi-source heterogeneous data to divide the data volume level. The data volume grading unit divides the data volume level of the multi-source heterogeneous data into three levels according to the size of the data volume; The data multi-dimensional classification module is further provided with a time sensitivity analysis unit, which is used to divide the delay tolerance level according to the source identifier. The time sensitivity analysis unit divides the delay tolerance level of the multi-source heterogeneous data into three levels according to different source identifiers; 4. The computer data processing system based on machine learning according to claim 3, characterized in that, The data multi-dimensional classification module is further provided with a three-dimensional classification construction unit, which is used to encode the delay tolerance level, data volume level, and security level to obtain a digital level, and quantitatively represent the delay tolerance level, data volume level, and security level; The three-dimensional classification construction unit analyzes the type judgment parameters according to the digital levels of the delay tolerance level, data volume level, and security level of the multi-source heterogeneous data to obtain the type judgment parameter S, and performs comparison analysis on the type judgment parameter to determine the comprehensive data type. The comprehensive data type of the multi-source heterogeneous data is divided into four categories according to the size of the type judgment parameter. The comprehensive data type includes emergency security affairs, emergency standard affairs, regular standard affairs, and batch processing affairs.
5. The computer data processing system based on machine learning according to claim 4, wherein The type feature analysis module performs feature analysis on the comprehensive data type, timestamp number, and multi-source heterogeneous data to obtain data feature L(d).
6. The computer data processing system based on machine learning according to claim 5, characterized in that, The priority scoring module is provided with a data scoring analysis unit, which is used to analyze the level scores according to the delay tolerance level, data volume level, and security level of the multi-source heterogeneous data to obtain a security level score B1, a data volume level score B2, and a delay tolerance level B3; The priority scoring module is also provided with a scoring weighted analysis unit, which is used to perform weighted analysis on the level score and data characteristics to obtain the priority score. The priority score is set as Z. If L(d) > θ, set Z = β1×B1 + β2×B2 + β3×B3. Conversely, set: Z = [β1 + L(d) / 10]×B1 + [β2 - L(d) / 5]×B2 + [β3 + L(d) / 10]×B3. Where, β1 represents the first weight, β2 represents the second weight, β3 represents the third weight, β1 + β2 + β3 = 1 and β1×β2×β3 ≠ 0.
7. The computer data processing system based on machine learning according to claim 6, wherein The priority scoring module is also provided with a classification matching and updating unit, which is used to perform matching analysis on the priority score and the data comprehensive type, extract the priority scores corresponding to different data comprehensive types as the comprehensive type scores, and respectively arrange the comprehensive type scores corresponding to each data comprehensive type in ascending order, and analyze the type discrete parameter according to the arranged comprehensive type scores to obtain the type discrete parameter R(m). The classification matching and updating unit performs comparison and analysis on the type discrete parameters to update the three-dimensional classification model. When R(m) > r, priority labels are added to the multi-source heterogeneous data that satisfies Z ≤ Z 0.5×NU(m) + Min(Z ∈ U(m)), where r represents the update matching threshold.
8. The computer data processing system based on machine learning according to claim 2, wherein The multi-source data processing module responds to and processes multi-source heterogeneous data according to the data comprehensive type and the priority score order. The multi-source data processing module respectively sorts the multi-source heterogeneous data with and without priority tags of the same data comprehensive type in descending order of the priority score, and preferentially processes the multi-source heterogeneous data with priority tags in the sorted order, and then processes the multi-source heterogeneous data without priority tags in the sorted order. If Z > z1, the multi-source data processing module allocates the multi-source heterogeneous data to the GPU acceleration queue for processing. If Z < z2, the multi-source data processing module classifies the multi-source heterogeneous data into the night batch task pool for processing. If z2 ≤ Z ≤ z1, the multi-source data processing module processes the multi-source heterogeneous data normally; where, z1 represents the first processing threshold, and z2 represents the second processing threshold.
9. The computer data processing system based on machine learning according to claim 1, characterized in that, The processing exception monitoring module is provided with a timeout processing judgment unit, which is used to judge the completion type of the multi-source heterogeneous data according to the processing duration, the timestamp number, and the delay tolerance level. When the delay tolerance level is the first level, if T ≤ γ1, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is normal completion; if T > γ1, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is timeout completion. If the processing duration is a missing value, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is uncompleted; when the delay tolerance level is the second level, if T ≤ γ2, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is normal completion. If T > γ2, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is timeout completion. If the processing duration is a missing value, the timeout processing judgment unit determines that the completion type of the multi-source heterogeneous data is uncompleted; when the delay tolerance level is the third level, the timeout processing judgment unit does not judge the completion type; where, T represents the processing duration, γ1 represents the first delay threshold, and γ2 represents the second delay threshold.
10. The computer data processing system based on machine learning according to claim 9, wherein The processing exception monitoring module is also provided with an exception handling analysis unit, which is used to count the number of times with the completion type of timeout completion as the timeout handling times, count the number of times with the completion type of uncompleted as the uncompleted times, and use the sum of 70% of the timeout handling times and the uncompleted times as the exception handling times; When the number of timeout handling times of multi-source heterogeneous data of the same data integration type in one second by the exception handling analysis unit is greater than or equal to k times, the rule optimization of the three-dimensional classification model is triggered. The rule optimization of the three-dimensional classification model is a weighted analysis process for optimizing the level scores according to the number of timeout handling times of multi-source heterogeneous data of the same data integration type in one second, optimizing the weights of each level score, correlating the weights of each level score with the logarithm of the timeout handling times, so as to optimize the three-dimensional classification model.
Citation Information
Patent Citations
Electric power heterogeneous data processing method based on edge computing
CN112130999A
Multi-source heterogeneous data cross-system cooperative processing method and system
CN114756883A
Data processing method and device based on multi-source heterogeneous information
CN118467937A
Forestry information sharing method and system based on multi-source data
CN118797530A
Multi-system fusion interaction method and system based on power protection platform
CN118797555A