Data processing method and system based on standardized information flow

By using a data processing method based on standardized information flow, the problem of mismatch between data processing and business scenarios caused by the dispersed data sources of enterprises is solved. It realizes accurate standardization and adaptability processing of multimodal data, and improves data processing efficiency and matching degree with business scenarios.

CN121579856APending Publication Date: 2026-02-27QINGDAO METRO GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511676252.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

In existing technologies, enterprise data suffers from scattered sources, resulting in chaotic feature identification metadata, low matching degree between data processing results and business scenarios, low efficiency of cross-system data integration, lack of clear classification in the management of derived features, redundancy in feature groups, and inability of data encapsulation to dynamically adapt to scenario requirements, leading to a mismatch between the data processing process and business scenarios.

Method used

By employing data processing methods based on standardized information flow, including metadata verification, classification feature decomposition, real-time and predictive correlation, derived feature classification and verification, and adaptive scene output encapsulation, we can achieve accurate standardization and adaptability processing of multimodal data.

Benefits of technology

It improves the accuracy and standardization of data in specific scenarios, ensures that data performance is adapted to scenario requirements, improves data processing efficiency and consistency, enhances business operation efficiency and data adaptability, meets the business scenario requirements of instant response and high processing efficiency, and improves the accuracy of scenario prediction and the adaptability of data encapsulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121579856A_ABST
    Figure CN121579856A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and system based on a standardized information flow, and relates to the technical field of electrical digital technology processing. According to the method, enterprise data processing is carried out through standardized features based on multi-modal enterprise data obtained after preprocessing, meanwhile, enterprise data association is carried out according to an enterprise data processing result, then derivative feature classification is carried out based on the enterprise data obtained after enterprise data association, business scene simulation is carried out on derivative features, and the business scene simulation result is obtained. The method comprises the steps of classifying standardized enterprise data according to derivative features, performing dimension verification according to the standardized enterprise data after derivative feature classification, performing standardized output packaging and standard format labeling of a self-adaptive scene based on scene collaboration data obtained after dimension verification, outputting a current scene decision action, and performing data iteration updating after a decision is completed. The problem that the matching degree of the data processing process and the business scene is not high due to insufficient enterprise data standardization in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of digital data processing, and in particular to a data processing method and system based on standardized information flow. BACKGROUND

[0002] With the deepening of the digital economy era, enterprises combine their own multi-business track synergy, integrate data processing flow with business scenarios, and current enterprise data contains not only structured data generated by traditional business systems, but also time series data collected by Internet of Things devices, unstructured data generated by Internet platforms, and semi-structured data provided. The data processing flow first relies on big data analysis technology to cover the format standard, semantic standard, and structure standard system from the enterprise historical mass data and business scenario demand. In the data preprocessing link, relying on the distributed architecture of big data, the parallel processing of multi-source data is realized. In the data integration stage, the text data is extracted to the enterprise feature structure and the structured data is standardized. Data enhancement generates derived features through logistic regression and collaborative filtering to enrich data. Finally, through adaptive packaging technology, the standardized data is adapted according to the business scenario demand, so that each step has historical enterprise data and corresponding scene for reference and strategy.

[0003] For example, the Chinese invention patent with publication number CN120336307B discloses a big data standardization method and system based on data integration, which includes: through a dynamic format adapter, based on data conversion mapping rules combined with reinforcement learning Schema algorithm, the original data of different data sources is converted into intermediate format data, the quantum superposition state of each field is obtained by using quantum state coding, the best mapping node in the knowledge graph is solved based on quantum annealing algorithm, and finally the standardization data is obtained by standardizing and mapping the solved field according to the best mapping node combined with the semantic constraints and standard format of the node.

[0004] For example, the Chinese invention patent application with publication number CN115687327A discloses a method of assisting data standardization by using big data, which includes: collecting data in the production management process and the warehouse management process, selecting name data in the same period to group by material category, and classifying in detail according to the classification rules of the same material, confirming the rationality of the establishment of standard name in the classification process, writing the name relationship into the standard data mapping relationship table after determination, and dividing the remaining materials into small categories, and finally collecting historical data for data cleaning.

[0005] The above-mentioned technology at least has the following technical problems:

[0006] In the prior art, enterprise data is sourced from multiple channels such as internal, external and third-party cooperation, and the field naming and data types of each channel are different. For example, in subway train monitoring, the data type corresponding to the arrival time of the train is a timestamp, while in ticketing monitoring, the data type corresponding to the entry time is a string. In the process of converting heterogeneous data into a unified format, enterprise data base features are generated based on general rules, which cannot filter core data according to scene requirements, resulting in a deviation between the processing result and the scene requirement. Different business scenarios have different performance requirements for processed data, such as real-time monitoring which requires immediate response, and scene maintenance which requires high processing efficiency for batch processing of enterprise data. In the process of enterprise information retrieval, the lack of enterprise data standardization results in a low matching degree between the data processing process and the business scenario. SUMMARY

[0007] To solve the technical problem of the prior art that the lack of enterprise data standardization results in a low matching degree between the data processing process and the business scenario, embodiments of the present application provide a data processing method and system based on standardized information flow. The technical solution is as follows:

[0008] On the one hand, a data processing method based on standardized information flow is provided, which includes the following steps: step one, based on the multi-modal enterprise data obtained after preprocessing, enterprise data processing is performed through standardized features to improve business operation efficiency and adaptability, and enterprise data correlation is performed according to the results of enterprise data processing to improve the coherence of enterprise data; step two, based on the enterprise data after enterprise data correlation, derivative feature classification is performed to simulate business scenarios for derivative features, and dimension verification is performed on the standardized enterprise data after derivative feature classification to ensure the eligibility of standardized enterprise data; step three, based on the scene collaborative data obtained after dimension verification, standardized output packaging and standard format labeling of adaptive scenes are performed to obtain enterprise processing data, while outputting the current scene decision action, and performing data iteration update after decision completion to improve the matching degree between enterprise processing data and business scene requirements.

[0009] In another aspect, a data processing system based on standardized information flow is provided, which is applied to a data processing method based on standardized information flow, and the system comprises an enterprise data association module, a derived feature classification dimension verification module and a standard format packaging update module; the enterprise data association module is used for processing enterprise data based on multi-modal enterprise data obtained after preprocessing through standardized features, and simultaneously performing enterprise data association according to the results of enterprise data processing; the derived feature classification dimension verification module is used for classifying derived features based on enterprise data after enterprise data association, simulating business scenarios for the derived features, and performing dimension verification according to standardized enterprise after derived feature classification; and the standard format packaging update module is used for obtaining scene coordination data after dimension verification, performing standardized output packaging and standard format labeling of adaptive scenes, simultaneously outputting current scene decision actions and performing data iterative update after decision completion.

[0010] The technical scheme provided by the embodiment of the application has at least the following beneficial effects:

[0011] 1. Through the meta-information verification and differential processing steps, the scene-based accurate standardization of multi-modal data is realized. In the prior art, multi-modal enterprise data often has chaotic feature identification meta-information (such as inconsistent format and missing key information) due to scattered sources (such as device sensing, business systems and text logs), and subsequent processing is prone to data deviation. First, a scene network topology structure is constructed based on a specified business scenario, and meta-information consistency verification is performed on the feature identification of multi-modal enterprise data. This verification link can accurately judge the accuracy and integrity of the feature identification basic information, and avoid deviation caused by basic information errors in subsequent processing. For data meeting the meta-information consistency, through classification feature disassembly, the feature field reading time and analysis time are analyzed in combination with the current business scenario, and the data reading and analysis efficiency is improved through specific means such as modifying the number of fragments, adjusting the number of parallel threads and optimizing the cache validity period, so as to ensure that the data performance adapts to the scene demand, provide high-quality standardized data basis for subsequent enterprise data association and decision-making, and improve the business operation efficiency and data adaptability.

[0012] 2. Due to the low efficiency of cross-system enterprise data integration, poor format compatibility, and lack of trend prediction capabilities, this solution divides enterprise data association into real-time association and predictive association. For real-time association, it extracts association keys from multi-source real-time data based on enterprise data processing results. Utilizing refined real-time association logic, it can quickly integrate multi-source real-time data, solving the problems of low efficiency and format incompatibility in existing cross-system data integration technologies, and meeting the immediate decision-making needs of scenarios such as real-time monitoring and real-time risk control. For predictive association, it obtains potential association sequences through historical association keys and adjusts the time period length based on time series deviations. The synergy between real-time and predictive methods enables enterprise data to achieve instant linkage and trend association with business data, improving the consistency of enterprise data and providing comprehensive data support for the entire business data processing process.

[0013] 3. Due to the lack of clear classification in the management of derived features, data drift (such as the shift in feature distribution over time) is not handled in a timely manner, leading to a decline in feature effectiveness. This solution achieves precise management and stable quality of derived features through classification allocation and dynamic correction. First, based on the enterprise data association results and feature identifiers, derived features are divided into two categories: homogeneous derived features and heterogeneous derived features. For homogeneous derived features, the impact of data drift on feature quality is reduced from the source, avoiding the problem of decreased feature effectiveness due to data drift. For heterogeneous derived feature groups, the consistency of business logic among different types of derived features is verified, resolving the problem of prediction confusion caused by logical contradictions between different types of features. At the same time, scenario prediction deviations are detected in this process. Both feature conflict and data drift issues are resolved through classification allocation, and feature quality is kept in sync with scenario prediction requirements through real-time monitoring and correction, improving the accuracy of scenario prediction and providing reliable feature support for business decisions.

[0014] 4. In existing technologies, feature groups often contain a large number of redundant features, leading to wasted computing resources, decreased prediction efficiency, and a lack of effective verification of inter-group collaboration. This solution achieves comprehensive protection of derived features through intra-group and inter-group collaboration verification. For intra-group collaboration verification, redundant features are effectively eliminated based on feature similarity within the group, avoiding the problem of wasted computing resources and decreased prediction efficiency caused by feature redundancy. For inter-group collaboration verification, a comparison logic between combined accuracy and intra-group single accuracy is introduced. Intra-group single accuracy reflects the accuracy of independent prediction by a single sub-feature group, while combined accuracy reflects the accuracy of joint prediction by different derived feature groups. By comparing the two, it is possible to intuitively determine whether inter-group collaboration generates positive value, solving the problem of intra-group feature redundancy and ensuring the effectiveness of inter-group collaboration. This allows derived feature groups to remain lightweight while achieving efficient collaboration, further improving the efficiency and accuracy of scene prediction.

[0015] 5. Existing technologies often use fixed formats and parameters for data encapsulation, making it difficult to dynamically adjust compression ratios and the number of blocks to adapt to different scenarios and performance requirements. This solution addresses this by using scenario awareness and dynamic adaptation to achieve standardized output encapsulation that is compatible with business scenarios. First, based on the enterprise data acquisition and processing scenario identifier after dimensional verification, the solution determines the appropriate decision-making scenario by combining it with a preset scenario requirement matrix and matching it with the scenario network topology to ensure that the encapsulation direction aligns with the core requirements of the scenario. During the encapsulation parameter adjustment phase, the compression ratio of the data compression engine and the number of data blocks per partition are dynamically adjusted to achieve precise adaptation. After encapsulation, consistency verification is performed to ensure that the standardized encapsulation package meets the performance requirements of different scenarios while ensuring format compliance and data quality, providing efficient and reliable data support for business scenario decisions. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart of a data processing method based on standardized information flow provided in an embodiment of the present invention;

[0018] Figure 2 A flowchart corresponding to enterprise data association provided in an embodiment of the present invention;

[0019] Figure 3 A flowchart illustrating the corresponding process of derived feature classification dimension verification and standard format encapsulation update provided in this embodiment of the invention;

[0020] Figure 4 This is a schematic diagram of the structure of a data processing system based on standardized information flow provided in an embodiment of the present invention. Detailed Implementation

[0021] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0022] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0023] In the embodiments of the present application, sometimes the subscript such as W1 may be written in the form of non-subscript such as W1, and when the difference is not emphasized, the meanings expressed are consistent.

[0024] To make the technical problems, technical solutions and advantages to be solved by the present application clearer, specific embodiments will be described in detail below with reference to the drawings.

[0025] The embodiments of the present application provide a data processing method based on standardized information flow, which can be implemented based on a data processing system of standardized information flow. Figure 1 As shown in the flowchart of the data processing method based on standardized information flow, the processing flow of the method can include the following steps: step one, based on the multi-modal enterprise data obtained after preprocessing, enterprise data processing is performed through standardized features to improve business operation efficiency and adaptability, and enterprise data correlation is performed according to the results of enterprise data processing to improve the coherence of enterprise data; step two, based on the enterprise data after enterprise data correlation, derivative feature classification is performed to simulate business scenarios for derivative features, and dimension verification is performed on the standardized enterprise data after derivative feature classification to ensure the eligibility of standardized enterprise data; step three, based on the scene collaborative data obtained after dimension verification, adaptive scene standardized output packaging and standard format labeling are performed to obtain enterprise processing data, and the current scene decision action is output, and data iteration update is performed after decision completion to improve the matching degree of enterprise processing data and business scenario demand.

[0026] In a specific embodiment, such as a subway group, the preprocessed multi-modal data (such as train operation monitoring data, gate passenger flow data, device sensor data, etc.) can be processed by standardized features to unify and adapt the data originally from different systems and in different formats, thereby improving the business operation efficiency in daily subway operation. At the same time, through enterprise data correlation, train operation status, passenger flow changes, device health data, etc. are connected to enhance the coherence between data and avoid operation decision lag caused by data fragmentation.

[0027] On this basis, the associated data is classified by derivative features, which can be simulated for different business scenarios such as morning and evening peak scheduling and device preventive maintenance, and then verified by dimensions to ensure data eligibility and provide reliable basis for subsequent decision-making. Finally, based on scene collaborative data, adaptive standardized output packaging and labeling are performed, which not only can accurately obtain enterprise processing data that meets the scene demand and output scene decision actions such as train departure adjustment and passenger flow diversion, but also can continuously optimize the matching degree of data and subway operation scenario demand through data iteration update after decision completion, thereby helping subway operation to be more efficient and decision-making to be more accurate.

[0028] As shown in the flowchart of the data processing method based on standardized information flow, the processing flow of the method can include the following steps: step one, based on the multi-modal enterprise data obtained after preprocessing, enterprise data processing is performed through standardized features to improve business operation efficiency and adaptability, and enterprise data correlation is performed according to the results of enterprise data processing to improve the coherence of enterprise data; step two, based on the enterprise data after enterprise data correlation, derivative feature classification is performed to simulate business scenarios for derivative features, and dimension verification is performed on the standardized enterprise data after derivative feature classification to ensure the eligibility of standardized enterprise data; step three, based on the scene collaborative data obtained after dimension verification, adaptive scene standardized output packaging and standard format labeling are performed to obtain enterprise processing data, and the current scene decision action is output, and data iteration update is performed after decision completion to improve the matching degree of enterprise processing data and business scenario demand. Figure 2As shown, the process diagram corresponding to the enterprise data association provided by the embodiment of the application is obtained by standardizing feature alignment for enterprise data association, wherein the standardized feature alignment is divided into classification feature disassembly and topology format alignment, the classification feature disassembly classifies the fields into first type fields, second type fields and third type fields and has different processing methods, the topology format alignment is to judge whether the constraint condition is matched first, if matched, the node is updated, if not matched, the conflict detection is performed; and the enterprise data association includes real-time association and predictive association.

[0029] Further, the enterprise data is processed by standardized features, and the specific process is: based on a specified business scenario, a scene network topology structure is obtained, and the feature identification of the multi-modal enterprise data is subjected to meta information consistency rate verification; if the meta information consistency rate corresponding to the obtained feature identification is not less than the preset meta information consistency rate, the enterprise data is processed, otherwise the meta information consistency rate verification is performed again, the meta information consistency rate is used to reflect the completeness of the feature identification process, and is obtained by the ratio of the number of obtained feature identifications to the total number of feature identifications; if the meta information consistency rate corresponding to the feature identification after re-verification is still less than the preset meta information consistency rate, it indicates that the enterprise data feature identification is missing, and an identification warning is prompted; based on the constructed scene network topology structure, the obtained multi-modal enterprise data is subjected to standardized feature alignment, and the specific process is: for the enterprise data whose meta information consistency rate corresponding to the feature identification is not less than the preset meta information consistency rate, classification feature disassembly is performed to improve the data reading and parsing efficiency; for the enterprise data whose meta information consistency rate corresponding to the feature identification is less than the preset meta information consistency rate, it is recorded as no identification data and subjected to topology format alignment to improve the business adaptability.

[0030] Specifically, the classification feature disassembly specific process is: through the acquired feature field reading time and combined with the current business scenario, scene identification classification is carried out, the feature field reading time represents the total time consumed from initiating the feature field reading request to successfully acquiring the field data; if the acquired feature field reading time is greater than the maximum value of the preset reading time interval, it is recorded as the first type field, indicating that this type of field is a to-be-assigned field, and the number of feature field fragments is corrected according to the reading time deviation to improve the parsing efficiency; if the acquired feature field reading time is within the preset reading time interval, the feature field parsing time is acquired, if the feature field parsing time is greater than the preset parsing time, it is recorded as the second type field, indicating that this type of field is a to-be-optimized field, and the number of loaded parallel threads is adjusted according to the parsing time deviation, the parallel data reading efficiency is improved to shorten the parsing time, if it is not greater than the preset parsing time, it is classified as the third type field; at the same time, for the second type field, if the feature field reading frequency is greater than the preset reading frequency, the field filtering strength is obtained according to the current acquired reading frequency to reduce the field proportion of invalid fields in the feature reading process, if not, it is classified as the third type field, the feature field parsing time represents the time consumed for format conversion, verification, structured processing and other parsing operations on the acquired feature field data, the feature field reading frequency represents the number of reading requests for the field per unit time; if the acquired feature field reading time is less than the minimum value of the preset reading time interval, it is recorded as the third type field, the reading efficiency of this type of field is qualified; the processing priority of the first type field, the second type field and the third type field decreases in turn.

[0031] Next, the topology format is aligned, and the specific steps are: the acquired unidentified data is matched with the preset topology constraint condition to determine the element node in the scene network topology structure, the preset topology constraint condition represents the result derived based on the historical operation law of the scene network topology structure and the interaction logic of the element node; if the unidentified data does not match the preset topology constraint condition, the preset personnel is prompted to perform format conflict detection, otherwise the unidentified data is updated as a new element node of the corresponding scene network topology structure; enterprise data association includes real-time association for quantifying the real-time matching degree and linkage efficiency of dynamic data in enterprise business processes, and predictive association for quantifying the association strength between historical data and current trends; and the specific process of real-time association is: based on the enterprise data processed by the enterprise data, the association key under the corresponding scene of the multi-source real-time data is acquired, and the association key is analyzed for data field matching, the association key represents the field combination of the same business scene object in the multi-modal enterprise data.

[0032] The data field matching analysis is specifically: if the data field corresponding to the association key matches more than a preset matching number of data fields, the corresponding data field is sorted according to a set business priority, the data field with the highest business priority is associated with the association key, and if there is more than one data field with the highest business priority, a preset secondary sorting rule is further used for screening; if the data field corresponding to the association key does not match a data field, real-time data is checked and supplemented, specifically: based on the data source response time length deviation, the number of asynchronous verification nodes is adjusted, and the number of verification nodes is increased to ensure the accuracy of the alignment of the association key and the time sequence in the association data pool; if the data field corresponding to the association key matches no more than a preset matching number of data fields, the current corresponding matched data is associated; the specific process of predictive association is: obtaining a potential association sequence corresponding to a preset business association scene in a historical association key; if the time sequence deviation corresponding to the potential association sequence is less than a preset time sequence deviation, the length of the read time period is processed based on the current obtained time sequence deviation, so as to align the potential association sequence with the business cycle rule, otherwise the association sequence corresponding to the predictive association is applied to the corresponding potential association sequence, and the time sequence deviation represents the time difference value between the current predicted business association node and the actual business association node.

[0033] In the embodiment, based on the specified business scene, the core business objects (such as devices, process nodes, data sources, etc.) in the specified business scene are combed as element nodes, and then the interaction rules and association relationships between the element nodes are determined according to the business logic, and a scene network topology structure is formed. The preset meta-information consistency rate is pre-set, usually 90%, the preset read time length interval is a value interval obtained by analyzing historical read time length data, storage read-write speed, and considering that the read time length is concentrated in which interval as the basic reference in most cases, and the interval can be adjusted in actual application. The preset read frequency is the result of summing and averaging the read frequency in the historical classification feature disassembly process, the read time length deviation is the difference between the obtained read time length and the preset read time length, and the preset read time length is a time length pre-set according to the current business scene and meeting the scene application purpose.

[0034] The data slice management adjustment interface is called according to the reading duration deviation, the number of slices of the feature field is increased or decreased to optimize the data analysis efficiency; the field filtering strength is obtained according to the reading frequency, the difference value is input into the preset mapping set of reading frequency difference-field filtering strength for mapping matching to obtain the field filtering strength of the feature field, the field filtering strength is used to measure the degree of filtering of the feature field, so as to reduce the field proportion of invalid fields in the feature reading process, the preset matching number of data fields is usually 2, and the set business priority is based on the demand of big data analysis of each business scene, for example, the real-time demand is the highest when the current scene is train operation state monitoring, similarly, the preset secondary sorting rule is similar to the business priority, the data source response duration deviation is the difference between the data source response duration and the preset data source response duration, the preset data source response duration is represented by the result of summing and averaging the historical data source response duration, the increase in the number of nodes is obtained by inputting the deviation into the mapping set of response deviation-checking node for mapping matching to obtain the increase value of the number of asynchronous checking nodes, so as to increase the checking link, the preset time sequence deviation is to first determine the maximum time deviation (such as a smaller deviation is required for fault warning) that can be accepted by the business scene (such as subway equipment fault early warning and passenger flow peak prediction), then statistics the time sequence deviation distribution of the same kind of prediction in history, and take the reasonable value within the normal fluctuation, finally determine the preset value combined with the business accuracy target, and process the reading time period length, if the time sequence deviation is that the prediction node is earlier than the actual node (such as predicting the passenger flow peak at 8 o'clock, and the actual time is 8 minutes and 10 seconds), then the original reading time period length (10 minutes) is increased according to the deviation; if the time sequence deviation is that the prediction node is later than the actual node, such as predicting the equipment fault at 15 o'clock, and the actual time is 14 minutes and 50 seconds, then the original reading time period length (10 minutes) is shortened according to the deviation, and the historical business cycle rule is referred to in the adjustment process to ensure that the new cycle does not deviate from the inherent rhythm of the business.

[0035] In the optimization of standardized feature processing and enterprise data association, first, the dual guarantee of data foundation quality and processing efficiency is realized. Through meta-information consistency rate verification, data with complete feature identification and accurate basic information can be screened out in advance, avoiding rework caused by missing identification or information errors, and reducing data quality problems from the source to interfere with business processes. For verified data, different reading and analysis efficiency fields are optimized according to business scenarios, which can not only improve the processing speed of inefficient fields by adjusting the number of shards, parallel threads and other methods, but also reduce resource waste by filtering invalid fields and optimizing cache expiration. After the proportion of invalid fields is reduced, data reading does not need to load redundant information, and cache strategy optimization reduces resource consumption caused by repeated loading, so that the data processing link can more efficiently adapt to business scenario requirements, whether it is real-time response business or batch processing business, and can obtain stable and adaptive data flow support, solving the problem of uneven data quality and processing efficiency in traditional processing, and providing high-quality and high-availability data foundation for subsequent business decision-making.

[0036] Secondly, the optimization of data association link improves the coherence and business adaptability of data, so that multi-source data can truly serve the whole-process business decision-making. Real-time association breaks down the barriers between different system data through association key matching and dynamic data checking and supplementing, ensuring that multi-source real-time data can be integrated immediately, avoiding business response delay caused by data gaps. For example, when the business needs multi-dimensional data linkage, it can quickly match the effective data that meets the demand, and through the adjustment of verification nodes to ensure time sequence alignment, further improve the accuracy of data matching. Predictive association adjusts the time sequence deviation to make the potential association sequence derived from historical data conform to the current business cycle rule, avoiding the prediction deviation caused by relying solely on historical data, so that data can not only support immediate decision-making for current business, but also provide reliable basis for future business trend prediction. Whether it is to respond to equipment state changes in advance or to reasonably plan business resource allocation, it can make more realistic judgments based on coherent and adaptive data, solving the problem of insufficient immediacy and large prediction deviation in traditional data association, and fully releasing the value of data in the whole business process.

[0037] As shown in Figure 3 The process diagram corresponding to the derived feature classification dimension verification and standard format packaging update provided by the embodiment of the application is shown in the figure, which divides the derived feature classification into the allocation of the same derived group and the different derived group, then judges whether the deviation is within the interval to perform filtering intensity correction and dimension verification, and the dimension verification includes inter-group collaborative verification and intra-group collaborative verification, after verification, dynamic adaptation and consistency verification are performed, according to the obtained data processing efficiency index to determine which interval it is in, to perform iterative update, cache duration adjustment and data matching warning.

[0038] Further, the derived feature classification is used to improve the accuracy of scene prediction. The specific process is: based on the associated enterprise data, the derived features are adjusted. The specific process is: based on the obtained feature identifier, the derived features are classified into: same derived feature allocation for overall correction of derived features of the same type, and different derived feature allocation for adaptation of derived features of different types. The same derived feature allocation is: based on the data drift rate deviation corresponding to all derived features in the same derived feature group, the time window length of the same derived feature group is corrected to reduce the influence of data drift. The different derived feature allocation is: the preset personnel detects the conflict of the different derived feature group, and after the conflict verification, the corresponding different derived feature group is allocated to the same derived feature. The conflict detection is: the preset personnel verifies whether the sub-feature group in the different derived feature group can be adapted to the same derived feature group. In the process of derived feature classification, the scene prediction deviation is monitored in real time. When the scene prediction deviation is within the preset scene prediction interval, the dimension verification is performed. Otherwise, based on the current obtained scene prediction deviation feature filtering strength, the correction is performed to improve the scene prediction integrity. The scene prediction deviation represents the deviation rate of the prediction result obtained by the generated derived feature through time series prediction from the reference result. The deviation rate is obtained by collecting historical derived feature time series data related to the current business scene (such as derived features in the past multiple periods), using the autoregressive integrated moving average model to output the prediction result, and the deviation degree between the prediction result and the pre-set reference result. The matching rate is similar to the processing method of the deviation rate. Scene prediction is used to predict derived features by using time series features in the autoregressive integrated moving average model.

[0039] Specifically, the dimension verification includes an intra-group collaborative verification for quantifying the degree of feature redundancy within a same-derived feature group, and an inter-group collaborative verification for quantifying the degree of inter-group collaboration of different-derived feature groups; the specific process of the intra-group collaborative verification is: based on the obtained intra-group feature similarity, for a feature pair with an intra-group feature similarity greater than a preset intra-group feature similarity, the feature with the highest priority in the feature pair is retained, if the priorities are the same, further filtering is performed based on a preset secondary sorting rule, and at the same time, the consistency of the intra-group verification after retaining the features is modified, specifically whether the business meanings of the retained intra-group features are directed to the same decision target (such as serving equipment fault early warning, rather than mixing irrelevant passenger flow features), if yes, they are same-derived features, otherwise, intra-group feature removal warning is performed; the specific process of the inter-group collaborative verification is: obtaining a combined accuracy rate and an intra-group single accuracy rate, if the combined accuracy rate is greater than the intra-group single accuracy rate, it indicates that the inter-group collaboration is qualified, otherwise, a combined warning is performed; the intra-group single accuracy rate represents the matching rate of the result obtained by taking the corresponding sub-feature group in the different-derived feature group as the input of scene prediction and the reference result, usually according to the application purpose of the current business scene, the sub-feature group with the highest priority is selected, such as the current scene has high real-time requirement, the time series sub-feature group is usually selected as the sub-feature group, and the combined accuracy rate represents the matching rate of the result obtained by taking the different-derived feature group as the input data of scene prediction and the reference result.

[0040] In this embodiment, the data drift rate deviation is the difference between the obtained data drift rate and the preset data drift rate. The preset data drift rate is obtained by collecting drift rate data of the feature group in the past period of time, analyzing the normal fluctuation range, taking the drift rate value that will not affect the prediction accuracy as the preset value, and updating the preset value by collecting new drift data. The time window correction value is obtained based on the data drift rate deviation as the core basis, combined with the business attributes of the feature group (such as the periodicity of the time sequence feature), and calibrated. First, the deviation, the original time window length, and the business attributes of the derived feature group (such as the subway passenger flow feature which needs to be matched with the morning and evening peak period) are input. If the drift rate deviation is positive, it means that the current window length needs to be shortened. The correction value is the product of the original window length minus the deviation value and the time window. If the deviation is negative, the current window correction value is the product of the original window length plus the deviation value and the time window. The preset scenario prediction interval is usually first based on the prediction deviation range acceptable by the business scenario, and then calibrated based on the normal fluctuation range of the historical similar prediction deviation. The filtering strength correction is obtained by the actual scenario prediction deviation minus the upper limit of the preset interval or the lower limit of the preset interval minus the actual scenario prediction deviation. The quantified out-of-interval deviation value is obtained as the core basis for strength correction. The preset deviation value-feature filtering strength adjustment amplitude mapping table (which is generated by historical scenario data calibration and has preset strength adjustment proportion corresponding to different out-of-interval deviation values) is called. If the out-of-interval deviation value is positive and larger, the larger the strength increase amplitude is determined according to the mapping table (such as the more the deviation value exceeds the upper limit, the higher the filtering strength increase proportion is, so as to eliminate more low-value features that interfere with the prediction). If the out-of-interval deviation value is negative and the absolute value is larger, the larger the strength decrease amplitude is determined according to the mapping table (such as the more the deviation value is below the lower limit, the higher the filtering strength decrease proportion is, so as to retain more features that may improve the prediction integrity).

[0041] Through the classification and dimension verification of derived features, the effects of scenario prediction and data quality guarantee are improved. In the classification of derived features, the adaptive adjustment of same and different derived features, combined with the dynamic correction of scenario prediction deviation, makes the derived features more accurately match the business scenario requirements, reduces the deviation of the prediction result from the actual result, and provides more reliable feature support for scenario simulation. In the dimension verification aspect, the in-group collaborative verification can effectively sort out the feature relationships within the same derived feature group, eliminate redundant features, avoid invalid information interference, and ensure the consistency of in-group verification, so that the synergy of features in the same group is more accurate.

[0042] The inter-group collaborative verification can clearly judge the collaborative effect between different derivative feature groups by comparing the combination and the intra-group single accuracy. If the collaboration is not good, it can be warned in time to ensure that different groups of features can produce better prediction value when combined. Overall, this series of operations not only optimizes the utilization efficiency of derivative features in scene prediction, but also strictly controls data quality from multiple dimensions of intra-group and inter-group to ensure that standardized enterprise data meets the verification requirements of business scenarios, laying a precise and reliable data foundation for subsequent business scenario simulation, decision output, etc. It reduces the prediction deviation or decision-making error caused by data dimension problems, and makes the data support for business scenarios more targeted and effective.

[0043] Further, the standardized output packaging and standard format labeling of adaptive scenes, the specific process is: based on the enterprise data after dimension verification, the processing data scene identifier is obtained, and the adaptive decision scene is obtained by combining the preset scene demand matrix, the adaptive decision scene is matched with the scene network topology structure, and dynamic adaptation is carried out based on the preset packaging format, the specific process is: through the scene adaptation component and combining the current scene time sequence priority, the compression ratio of the data compression engine is adjusted based on the transmission rate deviation, in order to reduce the interface response delay, the scene adaptation component represents a software or hardware module set with intelligent analysis and dynamic adjustment capability, which can deeply perceive various feature parameters of the current data processing scene, including but not limited to the real-time requirement of data, the complexity of business logic and the occupation situation of system resources, etc.; according to the efficiency priority of the current scene, and based on the transmission rate deviation, the number of single-partition data splitting blocks is adjusted to improve the scene query efficiency; after adjusting the scene plug-in, the packaging package is labeled based on the preset standard format, and the standard packaging package is obtained, and the consistency of the standard packaging package is verified to improve the standard packaging package format eligibility.

[0044] Specifically, the consistency check includes the following steps: based on the partition query duration corresponding to the partition encapsulation package and the partition data block size, a data processing efficiency index for quantifying the partition data processing efficiency is obtained; if the obtained data processing efficiency index is greater than the maximum value of the preset data processing efficiency index interval, it indicates that the standard encapsulation package and the data processing corresponding to the current specified scene are qualified, and data iteration update is performed; if the obtained data processing efficiency index is less than the minimum value of the preset data processing efficiency index interval, it indicates that the data processing is unqualified, and a data matching warning is sent; if the obtained data processing efficiency index is in the preset data processing efficiency index interval, the interface cache duration is adjusted based on the data processing efficiency index deviation; when the data processing efficiency index deviation is positive, the increase value of the cache duration is obtained, and the repetition rate of data query is reduced; when the data processing efficiency index is negative, the decrease value of the cache duration is obtained, so as to balance the cache efficiency and data real-time performance; the data efficiency index deviation is usually the difference between the middle value of the data processing efficiency index interval and the obtained data processing efficiency index; the data iteration update includes the following steps: based on the feedback of the standard encapsulation package result, the deviation set in the data processing process is adjusted and the business result of the adjusted encapsulation package is verified for incremental update; the deviation set represents the deviation data set obtained in the data processing process.

[0045] In the embodiment, the packaging format is SX-CP-001, indicating the first standard in the product standard sub-system in the target implementation standard system, wherein SX represents the target implementation standard system, CP represents the product standard sub-system, and 001 represents the serial number, which is used to indicate the position of the standard in the system. When a standard is added or deleted in the standard system, the serial number should be rearranged. The preset scene requirement matrix is a business scene requirement table constructed based on big data analysis. The scene time sequence priority is stored in the scene requirement matrix. The transmission rate deviation represents the difference between the obtained transmission rate and the preset transmission rate, and the preset transmission rate is preset according to historical transmission data combined with the current business scene requirement in the current adaptive scene model. The adaptive scene model stores a mapping set between the transmission rate deviation and the data compression ratio and the number of split blocks, and a mapping set between the data processing efficiency index and the interface cache time length. The transmission rate deviation, the current compression ratio and the number of split blocks, the data processing efficiency index and the current interface cache time length can be taken as input items, and the adjustment value can be taken as an output item. The adaptive scene model selects a lightweight regression model, collects historical data under multiple business scenes, and the dimensions include: business scene identification (such as real-time monitoring / batch analysis), historical actual transmission rate, preset transmission rate of the corresponding scene (calculate transmission rate deviation), historical adjusted compression ratio / split block number, and adjusted effect index (query time length / interface response delay) to design a loss function with the goal of adjusting the transmission rate to meet the scene requirements and optimize resource utilization. The model is obtained by iterative training of the training set. The preset data processing efficiency index interval is obtained by big data analysis, which is suitable for all business scenes, and is further adjusted according to the current scene requirements.

[0046] Specifically, the expression of the data processing efficiency index D is In the formula, S represents the actual data amount of each data block, T represents the actual partition query time length, log(N+1) represents the scheduling time loss caused by the increase of the number of blocks, N represents the number of blocks, and N represents the total number of data blocks split in the adaptive packaging link. The principle of the formula is to comprehensively measure the relationship between data size and time (including scheduling loss) to quantify the efficiency of partition data processing. The numerator represents the total data amount of the partition, representing the data size that needs to be processed, and the denominator represents the total time. The larger D is, the more effective data amount is processed in unit time, and the higher the data processing efficiency is. Conversely, the efficiency is lower.

[0047] In the steps of standardized output packaging and consistency checking of adaptive scenes, the optimization effect is concentrated in the dual improvement of packaging adaptability and data quality reliability, and can dynamically balance the data use efficiency and real-time demand. The adaptive packaging link perceives the characteristic parameters such as the timing priority and real-time requirement of the current scene through the scene adaptation component, and adjusts the data compression ratio and single-partition data split block quantity in combination with the transmission rate deviation; this makes the packaging package accurately fit the performance requirements of different scenes, such as reducing the interface delay by optimizing the compression ratio for real-time response type scenes, and improving the data reading efficiency by adjusting the split block quantity for batch query type scenes, avoiding the problem of difficult adaptation to multi-scene performance requirements, so that the standardized packaging package not only conforms to the format specification, but also efficiently supports the operation of the current business scene.

[0048] The consistency checking link constructs a clear quality judgment standard through quantitative data processing efficiency indicators (combined with query time score and block size), can accurately identify whether the standardized packaging package adapts to the data processing requirements of the current scene, ensures that the subsequent business decision is based on high-quality data and avoids low-quality data interference with the business process; and when the indicators are in a reasonable range, the cache can be extended to reduce repeated queries when the indicators are biased, and the cache can be shortened to ensure data real-time when the indicators are biased, achieving dynamic balance of cache efficiency and data freshness. This series of optimization makes the standardized output packaging scene adaptation and quality controllable, providing efficient, reliable and demand-oriented data support for subsequent business scene decision-making, reducing the business efficiency loss caused by packaging inadaptation and poor data quality.

[0049] As shown in Figure 3 The structure schematic diagram of the data processing system based on standardized information flow provided by the embodiment of the application, comprising the following modules: enterprise data association module, derived feature classification dimension checking module and standard format packaging update module; the enterprise data association module is used for processing enterprise data based on the standardized features and associating enterprise data based on the preprocessed multi-modal enterprise data; the derived feature classification dimension checking module is used for classifying derived features based on the enterprise data associated enterprise data, simulating business scenes for the derived features, and checking dimensions based on the standardized enterprise after the derived feature classification; the standard format packaging update module is used for adaptive scene standardized output packaging and standard format labeling based on the scene cooperative data obtained after the dimension checking, and outputs the current scene decision action and iteratively updates the data after the decision is completed.

[0050] In the embodiment, the enterprise data association module, the derived feature classification dimension checking module and the standard format packaging update module form a cooperative relationship with data flow and progressive support as the core, and jointly constitute a complete link of standardized information flow data processing.

[0051] Among them, the enterprise data association module as the starting link of data processing is responsible for the standardized feature processing and association of the preprocessed multi-modal enterprise data, and the associated enterprise data output is directly used as the core input of the derived feature classification dimension verification module; the derived feature classification dimension verification module carries out derived feature classification and business scenario simulation based on the associated data, and then carries out dimension verification to ensure that the data meets the scene requirements, and the scene collaborative data obtained after verification becomes the basis for processing of the standard format packaging update module; the standard format packaging update module completes the standardized output packaging and labeling of the adaptive scene relying on the scene collaborative data, and outputs the current scene decision action, and after the decision is completed, data iteration update is also carried out, which indirectly provides reference for the optimization of each module in the subsequent data processing process. The three modules are closely linked, the processing result of the previous module provides necessary data support for the next module, and finally the transformation from the initial multi-modal data to the enterprise processing data meeting the business requirements is realized.

[0052] The above embodiments can be implemented in whole or in part by software, hardware (such as a circuit), firmware, or any combination thereof. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the flow or function according to the embodiments of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid state disk.

[0053] It should be understood that the term "and / or" herein merely describes the association relationship of the associated objects, and indicates that there can be three relationships, for example, A and / or B can represent the three cases of A alone, A and B together, and B alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents that the front and rear associated objects are in an "or" relationship, but can also represent an "and / or" relationship, which can be understood according to the context.

[0054] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or similar expressions means any combination of the items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.

[0055] It should be understood that the size of the sequence number of the above-mentioned processes does not mean the order of execution in various embodiments of the present application. The execution order of the processes should be determined according to their functions and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0056] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0057] Those of ordinary skill in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-mentioned devices, apparatuses and units can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.

[0058] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0059] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e. they can be located in one place or distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.

[0060] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.

[0061] If the function is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the part of the present application that essentially contributes to the prior art or the part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0062] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A data processing method based on standardized information flow, characterized in that, The method comprises the following steps: Step one, based on the multi-modal enterprise data obtained after preprocessing, the enterprise data is processed through standardized features to improve business operation efficiency and adaptability, and the enterprise data is associated according to the results of enterprise data processing to improve the coherence of enterprise data; Step two, based on the enterprise data after association, the derived features are classified to simulate business scenarios, and the standardized enterprise data after classification of derived features is dimensionally verified to ensure the eligibility of standardized enterprise data; Step three, based on the scenario collaborative data obtained after dimension verification, the standardized output packaging and standard format labeling of adaptive scenarios are carried out to obtain enterprise processing data, and the current scene decision action is output, and the data is iteratively updated after the decision is completed to improve the matching degree of enterprise processing data and business scenario demand.

2. The data processing method based on the standardized information flow according to claim 1, wherein, The specific process of processing enterprise data through standardized features is as follows: Based on the specified business scenario, the scene network topology structure is obtained, and the meta information consistency rate of the feature identification of multi-modal enterprise data is verified: If the meta information consistency rate corresponding to the obtained feature identification is not less than the preset meta information consistency rate, the enterprise data is processed, otherwise the meta information consistency rate is verified again, and the meta information consistency rate is used to reflect the completeness of the feature identification process; If the meta information consistency rate corresponding to the re-verified feature identification is still less than the preset meta information consistency rate, it means that the enterprise data feature identification is missing, and an identification warning is prompted; The enterprise data processing is as follows: based on the constructed scene network topology structure, the obtained multi-modal enterprise data is aligned with the standardized features: For enterprise data with meta information consistency rate corresponding to feature identification not less than the preset meta information consistency rate, the classification feature is disassembled to improve the data reading and parsing efficiency; For enterprise data with meta information consistency rate corresponding to feature identification less than the preset meta information consistency rate, it is recorded as no identification data and the topology format is aligned to improve business adaptability.

3. The data processing method based on the standardized information flow according to claim 2, wherein, The specific process of classification feature disassembly is as follows: Through the obtained feature field reading time and combined with the current business scenario, scene identification classification is carried out, and the feature field reading time represents the total time consumed from initiating the feature field reading request to successfully obtaining the field data; If the obtained feature field reading time is greater than the maximum value of the preset reading time interval, it is recorded as a first type field, which means that this type of field is a to-be-assigned field, and the number of feature field shards is corrected according to the reading time deviation to improve the parsing efficiency; If the obtained feature field reading time is within the preset reading time interval, the feature field parsing time is obtained, if the feature field parsing time is greater than the preset parsing time, it is recorded as a second type field, which means that this type of field is a to-be-optimized field, and the number of loaded parallel threads is adjusted according to the parsing time deviation, the parallel data reading efficiency is improved to shorten the parsing time, and if it is not greater than the preset parsing time, it is classified as a third type field; Meanwhile, for the second type of field, if the feature field reading frequency is greater than the preset reading frequency, the field filtering strength is obtained according to the current obtained reading frequency, so as to reduce the field proportion of invalid fields in the process of feature reading, and if not, it is classified as a third type of field; If the obtained feature field reading duration is less than the minimum value of the preset reading duration interval, it is recorded as a third type of field, and the reading efficiency of this type of field is qualified; The processing priority of the first type of field, the second type of field and the third type of field decreases in turn.

4. The data processing method based on the standardized information flow according to claim 2, wherein, The topology format is aligned, and the specific steps are: The obtained unmarked data is matched with the preset topology constraint condition to determine the element node in the scene network topology structure, and the preset topology constraint condition represents the result derived based on the historical operation law of the scene network topology structure and the interaction logic of the element node; If the unmarked data does not match the preset topology constraint condition, the preset personnel is prompted to detect the format conflict, otherwise the unmarked data is updated to the topology element node and added to the corresponding scene network topology structure as a new element node; The enterprise data association includes real-time association for quantifying the instant matching degree and linkage efficiency of dynamic data in enterprise business processes, and predictive association for quantifying the association strength between historical data and current trends; The specific process of the real-time association is: Obtain the association key in the corresponding scene of the multi-source real-time data, and simultaneously perform data field matching analysis on the association key, and the association key represents the field combination of the same business scene object in the multi-modal enterprise data.

5. The data processing method based on the standardized information flow according to claim 4, wherein, The data field matching analysis is specifically: If the data field corresponding to the association key matches more than a preset number of data fields, the corresponding data field is sorted according to a set business priority, the data field with the highest business priority is associated with the association key, and if there is more than one data field with the highest business priority, a preset secondary sorting rule is further used for screening; If the data field corresponding to the association key does not match a data field, real-time data is supplemented, specifically: Based on the data source response duration deviation, the number of asynchronous verification nodes is adjusted, and the number of verification nodes is increased to ensure the accuracy of the time sequence alignment of the association key and the association data pool; If the data field corresponding to the association key matches no more than a preset number of data fields, the current corresponding matched data is associated; The specific process of the predictive association is to obtain the potential association sequence corresponding to the preset business association scene in the historical association key; If the time sequence deviation corresponding to the potential association sequence is less than a preset time sequence deviation, the length of the read time period is processed based on the current obtained time sequence deviation, so as to align the potential association sequence with the business cycle law, otherwise the association sequence corresponding to the predictive association is applied to the corresponding potential association sequence, and the time sequence deviation represents the time difference value between the current predicted business association node and the actual business association node.

6. The data processing method based on the standardized information flow according to claim 1, wherein, The derived feature classification is used to improve the accuracy of scene prediction, and the specific process is: Based on the obtained feature identifier, the derived features are classified into same-derived feature assignments for overall correction of derived features of the same type and different-derived feature assignments for adaptation of derived features of different types; The same-derived feature assignment specifically includes correcting the time window length of the same-derived feature group based on the data drift rate deviation corresponding to all derived features in the same-derived feature group to reduce the impact of data drift; The different-derived feature assignment specifically includes conflict detection of the different-derived feature group by a preset person, and same-derived feature assignment of the corresponding different-derived feature group after conflict detection; During the classification of derived features, the scene prediction deviation is monitored in real time. When the scene prediction deviation is within a preset scene prediction interval, dimension checking is performed, otherwise, the filtering strength of the currently obtained scene prediction deviation feature is corrected to improve the completeness of scene prediction. The scene prediction deviation represents the deviation rate of the prediction result based on the generated derived feature from the reference result.

7. The data processing method based on the standardized information flow according to claim 6, wherein, The dimension checking includes in-group collaborative checking for quantifying the degree of feature redundancy in the same-derived feature group, and inter-group collaborative checking for quantifying the degree of inter-group collaboration of the different-derived feature group; The specific process of the in-group collaborative checking is as follows: Based on the obtained in-group feature similarity, for a feature pair with in-group feature similarity greater than a preset in-group feature similarity, the feature with the highest priority in the feature pair is retained. If the priorities are the same, further filtering is performed based on a preset secondary sorting rule, and the consistency of the same-derived feature group corresponding to the retained feature pair is verified and corrected; The specific process of the inter-group collaborative checking is as follows: The combined accuracy rate and the in-group single accuracy rate are obtained. If the combined accuracy rate is greater than the in-group single accuracy rate, it indicates that the inter-group collaboration is qualified, otherwise, a combined warning is performed; The in-group single accuracy rate represents the matching rate of the result obtained by taking the corresponding sub-feature group in the different-derived feature group as the input of scene prediction and the reference result. The combined accuracy rate represents the matching rate of the result obtained by taking the different-derived feature group as the input data of scene prediction and the reference result.

8. The data processing method based on the standardized information flow according to claim 1, wherein, The specific process of the standardized output packaging and standard format labeling of the adaptive scene is as follows: Based on the enterprise data after dimension checking, the processing data scene identifier is obtained, and the adaptive decision scene is obtained by combining the preset scene demand matrix. The adaptive decision scene is matched with the scene network topology structure, and dynamic adaptation is performed based on the preset packaging format. The specific process is as follows: Through the scene adaptation component and combining the current scene time sequence priority, the compression ratio of the data compression engine is adjusted based on the transmission rate deviation to reduce the interface response delay; According to the efficiency priority of the current scene, and based on the transmission rate deviation, the number of single-partition data splitting blocks is adjusted to improve the scene query efficiency; After adjusting the scene plug-in, the packaging package is labeled based on the preset standard format to obtain a standard packaging package, and consistency checking is performed on the standard packaging package to improve the format qualification of the standard packaging package.

9. The data processing method based on the standardized information flow according to claim 8, wherein, The specific process of the consistency checking is as follows: Based on the partition query duration corresponding to the partition encapsulation package and the partition data block size, a data processing efficiency index for quantifying the partition data processing efficiency is obtained; If the obtained data processing efficiency index is greater than the preset maximum value of the data processing efficiency index interval, it indicates that the standard encapsulation package and the current specified scene correspond to qualified data processing, and data iteration update is performed; If the obtained data processing efficiency index is less than the preset minimum value of the data processing efficiency index interval, it indicates that the data processing is unqualified, and a data matching warning is sent; If the obtained data processing efficiency index is in the preset data processing efficiency index interval, the interface cache duration is adjusted based on the data processing efficiency index deviation to balance the cache efficiency and data real-time performance. The data iteration update has the following specific process: based on the feedback of the standard encapsulation package result, the deviation set in the data processing process is adjusted and the business result of the adjusted encapsulation package is verified for incremental update, and the deviation set represents the deviation data set obtained in the data processing process.

10. A system for applying the data processing method based on normalized information flow according to any one of claims 1 to 9, comprising: An enterprise data association module, a derived feature classification dimension verification module, and a standard format encapsulation update module; The enterprise data association module is used for processing enterprise data based on the preprocessed multi-modal enterprise data through standardized features, and associating enterprise data according to the results of enterprise data processing. The derived feature classification dimension verification module is used for classifying derived features based on the enterprise data after enterprise data association, simulating business scenarios for derived features, and verifying dimensions based on standardized enterprises after derived feature classification. The standard format encapsulation update module is used for standardized output encapsulation and standard format labeling of the scene cooperative data obtained after dimension verification, outputting the current scene decision action, and performing data iteration update after the decision is completed.

Citation Information

Patent Citations

  • Method for assisting data standardization by using big data

    CN115687327A

  • A big data standardization method and system based on data integration

    CN120336307B